PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

iProClass: an integrated database of protein family, function and structure information.

The iProClass database provides comprehensive, value-added descriptions of proteins and serves as a framework for data integration in a distributed networking environment. The protein information in iProClass includes family relationships as well as structural and functional classifications and features. The current version consists of about 830 000 non-redundant PIR-PSD, SWISS-PROT, and TrEMBL proteins organized with more than 36 000 PIR superfamilies, 145 000 families, 4000 domains, 1300 motifs and 550 000 FASTA similarity clusters. It provides rich links to over 50 database of protein sequences, families, functions and pathways, protein-protein interactions, post-translational modifications, protein expressions, structures and structural classifications, genes and genomes, ontologies, literature and taxonomy. Protein and superfamily summary reports present extensive annotation information and include membership statistics and graphical display of domains and motifs. iProClass employs an open and modular architecture for interoperability and scalability. It is implemented in the Oracle object-relational database system and is updated biweekly. The database is freely accessible from the web site at http://pir.georgetown.edu/iproclass/ and searchable by sequence or text string. The data integration in iProClass supports exploration of protein relationships. Such knowledge is fundamental to the understanding of protein evolution, structure and function and crucial to functional genomic and proteomic research.

Amino Acid Motifs↗

Thermodynamic databases for proteins and protein-nucleic acid interactions.

Thermodynamic data regarding proteins and their interactions are important for understanding the mechanisms of protein folding, protein stability, and molecular recognition. Although there are several structural databases available for proteins and their complexes with other molecules, databases for experimental thermodynamic data on protein stability and interactions are rather scarce. Thus, we have developed two electronically accessible thermodynamic databases. ProTherm, Thermodynamic Database for Proteins and Mutants, contains numerical data of several thermodynamic parameters of protein stability, experimental methods and conditions, along with structural, functional, and literature information. ProNIT, Thermodynamic Database for Protein-Nucleic Acid Interactions, contains thermodynamic data for protein-nucleic acid binding, experimental conditions, structural information of proteins, nucleic acids and the complex, and literature information. These data have been incorporated into 3DinSight, an integrated database for structure, function, and properties of biomolecules. A WWW interface allows users to search for data based on various conditions, with different display and sorting options, and to visualize molecular structures and their interactions. These thermodynamic databases, together with structural databases, help researchers gain insight into the relationship among structure, function, and thermodynamics of proteins and their interactions, and will become useful resources for studying proteins in the postgenomic era.

DNA↗

ProTherm, Thermodynamic Database for Proteins and Mutants: developments in version 3.0.

The current release of ProTherm, Thermodynamic Database for Proteins and Mutants, contains more than 10 000 numerical data (300% of the first version) of several thermodynamic parameters, experimental methods and conditions, reversibility of folding, details about the surrounding residues in space for all mutants, structural, functional and literature information. In the current version, we have added information about the source of each protein, identification codes for SWISS-PROT and Protein Information Resource and unique Protein Data Bank (PDB) code for proteins with relevant source. We have also provided additional options to search for data based on PDB code, number of states and reversibility. ProTherm is cross-linked with other sequence, structural, functional and literature databases, and the mutant sites and surrounding residues are automatically mapped on the structure. The ProTherm database is freely available at http://www.rtc.riken.go.jp/jouhou/protherm/protherm.html.

Animals↗

The RESID Database of protein structure modifications and the NRL-3D Sequence-Structure Database.

The RESID Database is a comprehensive collection of annotations and structures for protein post-translational modifications including N-terminal, C-terminal and peptide chain cross-link modifications. The RESID Database includes systematic and frequently observed alternate names, Chemical Abstracts Service registry numbers, atomic formulas and weights, enzyme activities, taxonomic range, keywords, literature citations with database cross-references, structural diagrams and molecular models. The NRL-3D Sequence-Structure Database is derived from the three-dimensional structure of proteins deposited with the Research Collaboratory for Structural Bioinformatics Protein Data Bank. The NRL-3D Database includes standardized and frequently observed alternate names, sources, keywords, literature citations, experimental conditions and searchable sequences from model coordinates. These databases are freely accessible through the National Cancer Institute-Frederick Advanced Biomedical Computing Center at these web sites: http://www. ncifcrf.gov/RESID, http://www.ncifcrf.gov/NRL-3D; or at these National Biomedical Research Foundation Protein Information Resource web sites: http://pir.georgetown.edu/pirwww/dbinfo/resid .html, http://pir.georgetown.edu/pirwww/dbinfo/nrl3d .html

Amino Acids↗

Availability of short amino acid sequences in proteins.

Much attention is being paid to protein databases as an important information source for proteome research. Although used extensively for similarity searches, protein databases themselves have not fully been characterized. In a systematic attempt to reveal protein-database characters that could contribute to revealing how protein chains are constructed, frequency distributions of all possible combinatorial sets of three, four, and five amino acids ("triplets," "quartets," and "pentats"; collectively called constituent sequences) have been examined in the nonredundant (nr) protein database, demonstrating the existence of nonrandom bias in their "availability" at the population level. Nonexistent short sequences of pentats were found that showed low availability in biological proteins against their expected probabilities of occurrence. Among them, six representative ones were successfully synthesized as peptides with reasonably high yields in a conventional Fmoc method, excluding the possibility that a putative physicochemical energy barrier in forming them could be a direct cause for the low availability. They were also expressed as soluble fusion proteins in a conventional Escherichia coli BL21Star(DE3) system with reasonably high yield, again excluding a possible difficulty in their biological synthesis. Together, these results suggest that information on three-dimensional structures and functions of proteins exists in the context of connections of short constituent sequences, and that proteins are composed of evolutionarily selected constituent sequences, which are reflected in their availability differences in the database. These results may have biological implications for protein structural studies.

Amino Acid Sequence↗

Fast comparison of a DNA sequence with a protein sequence database.

We describe a computer program, named DNA-Protein Search (DPS), for comparing a megabase DNA sequence with a protein sequence database. The DPS program addresses the problems of frameshifts and introns in the DNA sequence. The DPS program was used to compare each of the following sequences with the Swiss-Prot database: the 1.8-megabase sequence of the Haemophilus influenzae Rd genome, the 0.58-megabase sequence of the Mycoplasma genitalium genome, and the 0.56-megabase sequence of Saccharomyces cerevisiae chromosome VIII. The comparisons found new regions that are similar to protein sequences. The sensitivity of DPS was evaluated using as test data the known coding regions of the three DNA sequences. The results demonstrate that the DPS program is a useful tool for finding the coding regions of the DNA sequence. The DPS program uses an order of magnitude less computer memory and is several times faster than the BLASTX program.

Amino Acid Sequence↗

Construction of validated, non-redundant composite protein sequence databases.

A strategy has been developed for the construction of a validated, comprehensive composite protein sequence database. Entries are amalgamated from primary source data bases by a largely automated set of processes in which redundant and trivially different entries are eliminated. A modular approach has been adopted to allow scientific judgement to be used at each stage of database processing and amalgamation. Source databases are assigned a priority depending on the quality of sequence validation and commenting. Rejection of entries from the lower priority database, in each pairwise comparison of databases, is carried out according to optionally defined redundancy criteria based on sequence segment mismatches. Efficient algorithms for this methodology are embodied in the COMPO software system. COMPO has been applied for over 2 years in construction and regular updating of the OWL composite protein sequence database from the source databases NBRF-PIR, SWISS-PROT, a GenBank translation retrieved from the feature tables, NBRF-NEW, NEWAT86, PSD-KYOTO and the sequences contained in the Brookhaven protein structure databank. OWL is part of the ISIS integrated data resource of protein sequence and structure [Akrigg et al. (1988) Nature, 335, 745-746]. The modular nature of the integration process greatly facilitates the frequent updating of OWL following releases of the source databases. The extent of redundancy in these sources is revealed by the comparison process. The advantages of a robust composite database for sequence similarity searching and information retrieval are discussed.

Amino Acid Sequence↗

ProTherm and ProNIT: thermodynamic databases for proteins and protein-nucleic acid interactions.

ProTherm and ProNIT are two thermodynamic databases that contain experimentally determined thermodynamic parameters of protein stability and protein-nucleic acid interactions, respectively. The current versions of both the databases have considerably increased the total number of entries and enhanced search interface with added new fields, improved search, display and sorting options. As on September 2005, ProTherm release 5.0 contains 17,113 entries from 771 proteins, retrieved from 1497 scientific articles (approximately 20% increase in data from the previous version). ProNIT release 2.0 contains 4900 entries from 273 research articles, representing 158 proteins. Both databases can be queried using WWW interfaces. Both quick search and advanced search are provided on this web page to facilitate easy retrieval and display of the data from these databases. ProTherm is freely available online at http://gibk26.bse.kyutech.ac.jp/jouhou/Protherm/protherm.html and ProNIT at http://gibk26.bse.kyutech.ac.jp/jouhou/pronit/pronit.html.

DNA↗

Improving reproducibility and sensitivity in identifying human proteins by shotgun proteomics.

Identifying proteins in cell extracts by shotgun proteomics involves digesting the proteins, sequencing the resulting peptides by data-dependent mass spectrometry (MS/MS), and searching protein databases to identify the proteins from which the peptides are derived. Manual analysis and direct spectral comparison reveal that scores from two commonly used search programs (Sequest and Mascot) validate less than half of potentially identifiable MS/MS spectra (class positive) from shotgun analyses of the human erythroleukemia K562 cell line. Here we demonstrate increased sensitivity and accuracy using a focused search strategy along with a peptide sequence validation script that does not rely exclusively on XCorr or Mowse scores generated by Sequest or Mascot, but uses consensus between the search programs, along with chemical properties and scores describing the nature of the fragmentation spectrum (ion score and RSP). The approach yielded 4.2% false positive and 8% false negative frequencies in peptide assignments. The protein profile is then assembled from peptide assignments using a novel peptide-centric protein nomenclature that more accurately reports protein variants that contain identical peptide sequences. An Isoform Resolver algorithm ensures that the protein count is not inflated by variants in the protein database, eliminating approximately 25% of redundant proteins. Analysis of soluble proteins from a human K562 cells identified 5130 unique proteins, with approximately 100 false positive protein assignments.

Cell Line, Tumor↗

High-throughput peptide mass fingerprinting of soybean seed proteins: automated workflow and utility of UniGene expressed sequence tag databases for protein identification.

Identification of anonymous proteins from two-dimensional (2-D) gels by peptide mass fingerprinting is one area of proteomics that can greatly benefit from a simple, automated workflow to minimize sample contamination and facilitate high-throughput sample processing. In this investigation we outline a workflow employing robotic automation at each step subsequent to 2-D gel electrophoresis. As proof-of-concept, 96 protein spots from a 2-D gel were analyzed using this approach. Whole protein (1 mg) from mature, dry soybean (Glycine max [L.] Merr.) cv. Jefferson seed was resolved by high resolution 2-D gel electrophoresis. Approximately 150 proteins were observed after staining with Coomassie Blue. The rather low number of detected proteins was due to the fact that the dynamic range of protein expression was greater than 100-fold. The most abundant proteins were seed storage proteins which in total represented over 60% of soybean seed protein. Using peptide mass fingerprinting 44 protein spots were identified. Identification of soybean proteins was greatly aided by the use of annotated, contiguous Expressed Sequence Tag (EST) databases which are available for public access (UniGene, ftp.ncbi.nih.gov/repository/UniGene/). Searches were orders of magnitude faster when compared to searches of unannotated EST databases and resulted in a higher frequency of valid, high-scoring matches. Some abundant, non seed storage proteins identified in this investigation include an isoelectric series of sucrose binding proteins, alcohol dehydrogenase and seed maturation proteins. This survey of anonymous seed proteins will serve as the basis for future comparative analysis of seed-filling in soybean as well as comparisons with other soybean varieties.

Databases, Genetic↗

The CATH extended protein-family database: providing structural annotations for genome sequences.

An automatic sequence search and analysis protocol (DomainFinder) based on PSI-BLAST and IMPALA, and using conservative thresholds, has been developed for reliably integrating gene sequences from GenBank into their respective structural families within the CATH domain database (http://www.biochem.ucl.ac.uk/bsm/cath_new). DomainFinder assigns a new gene sequence to a CATH homologous superfamily provided that PSI-BLAST identifies a clear relationship to at least one other Protein Data Bank sequence within that superfamily. This has resulted in an expansion of the CATH protein family database (CATH-PFDB v1.6) from 19,563 domain structures to 176,597 domain sequences. A further 50,000 putative homologous relationships can be identified using less stringent cut-offs and these relationships are maintained within neighbour tables in the CATH Oracle database, pending further evidence of their suggested evolutionary relationship. Analysis of the CATH-PFDB has shown that only 15% of the sequence families are close enough to a known structure for reliable homology modeling. IMPALA/PSI-BLAST profiles have been generated for each of the sequence families in the expanded CATH-PFDB and a web server has been provided so that new sequences may be scanned against the profile library and be assigned to a structure and homologous superfamily.

Algorithms↗

Puzzle pieces defined: locating common packing units in tertiary protein contacts.

Puzzle pieces are defined as small packing units which make up the unique tertiary interactions in proteins. Anti-parallel and perpendicular helix-helix contacts were broken down into basic puzzle-piece pairs in order to study the traits of such contacts: their limited geometry, preferred residue involvement, residue conformation and other common constraints. These traits can then be used for continued comparison of other protein structures, improving models of and designing proteins de novo and, in time, predicting 3D structure from primary sequence. Results from a small (100 proteins) database of anti-parallel helix-helix contacts and from preliminary work on a large database (600 proteins) of perpendicular helix-helix contacts are presented.

Amino Acid Sequence↗

Repair-FunMap: a functional database of proteins of the DNA repair systems.

UNLABELLED: Repair-FunMap is a functional database of the DNA repair systems. This database contains not only the proteins directly involved in DNA repair, but also the proteins that interact with the DNA repair proteins. A protein interaction network associated with the human DNA repair processes was established according to the functional relationship between proteins in the database. This network represents the current knowledge on the intrinsic signaling pathways related to DNA repair. The Repair-FunMap could become an essential resource center for cancer research, providing clues to understanding the inter-relationship between proteins in the network, and to building scientific models of the DNA repair processes. AVAILABILITY: http://astro.temple.edu/~feng/Servers/BioinformaticServers.htm

Abstracting and Indexing↗

Rapid 3D protein structure database searching using information retrieval techniques.

MOTIVATION: As the sizes of three-dimensional (3D) protein structure databases are growing rapidly nowadays, exhaustive database searching, in which a 3D query structure is compared to each and every structure in the database, becomes inefficient. We propose a rapid 3D protein structure retrieval system named 'ProtDex2', in which we adopt the techniques used in information retrieval systems in order to perform rapid database searching without having access to every 3D structure in the database. The retrieval process is based on the inverted-file index constructed on the feature vectors of the relationships between the secondary structure elements (SSEs) of all the 3D protein structures in the database. ProtDex2 is a significant improvement, both in terms of speed and accuracy, upon its predecessor system, ProtDex. RESULTS: The experimental results show that ProtDex2 is very much faster than two well-known protein structure comparison methods, DALI and CE, yet not sacrificing on the accuracy of the comparison. When comparing with a similar SSE-based method, namely TopScan, ProtDex2 is much faster with comparable degree of accuracy. AVAILABILITY: The software is available at: http://xena1.ddns.comp.nus.edu.sg/~genesis/PD2.htm

Algorithms↗

A contact scoring matrix for qualitative prediction of change in folding of alpha-helices in globular proteins caused by a mutation.

The atomic pairs in contact for atoms from pairs of amino-acid residues on pairs of helices in a protein database consisting of 48 proteins of known tertiary structure from the Brookhaven Protein Data Bank are searched and counted to construct a primary scoring system. Each score in the primary scoring system is weighted further with the possibility of occurrence of each residue pair in the protein database to give a final scoring matrix. Scores for predicting change in folding of alpha-helices in a mutant protein are calculated by assuming that every pair of helices in the protein can closely interact with each other. It is shown that the change in folding of alpha-helices in several mutant proteins are reflected in both the change of the contact scores and the helix geometry calculated.

Databases, Factual↗

PIDD: database for Protein Inter-atomic Distance Distributions.

Protein Inter-atomic Distance Distributions (PIDD) is a dedicated database and structural bio-informatics system for distance based protein modeling. The database is developed to host and analyze the statistical data for protein inter-atomic distances based on their distributions in databases of known protein structures such as in the Protein Data Bank (PDB). PIDD is capable of generating, caching, and displaying the statistical distributions of the distances of various types and ranges. The collected information can be used to extract geometric restraints or mean-force potentials for protein structure determination including nuclear magnetic resonance structure determination and comparative model refinement. PIDD is supported with a friendly designed web interface so that users can easily specify the distance types and ranges, and retrieve, visualize or download the distributions of the distances as they desire. PIDD is freely accessible at http://www.math.iastate.edu/pidd.

Databases, Protein↗

Update of KEYnet: a gene and protein names database for biosequences functional organisation.

KEYnet is a database where gene and protein names are hierarchically structured. Particular care has been devoted to the search and organisation of synonyms. The structuring is based on biological criteria in order to assist the user in data search and to minimise the risk of information loss. Links to the EMBL data library by the entry name and the accession number are implemented. KEYnet is available through the WWW at the following site: http://www.ba.cnr.it/keynet.html

Databases, Factual↗

The protein disease database of human body fluids: I. Rationale for the development of this database.

We are developing a relational database to facilitate quantitative and qualitative comparisons of proteins in human body fluids in normal and disease states. For decades researchers and clinicians have been studying proteins in body fluids such as serum, plasma, cerebrospinal fluid and urine. Currently, most clinicians evaluate only a few specific proteins in a body fluid such as plasma when they suspect that a patient has a disease. Now, however, high resolution two-dimensional protein electrophoresis allows the simultaneous evaluation of 1,500 to 3,000 proteins in complex solutions, such as the body fluids. This and other high resolution methods have encouraged us to collect the clinical data for the body fluid proteins into an easily accessed database. For this reason, it has been constructed on the Internet World Wide Web (WWW) under the title Protein Disease Database (PDD). In addition, this database will provide a linkage between the disease-associated protein alterations and images of the appropriate proteins on high-resolution electrophoretic gels of the body fluids. This effort requires the normalization of data to account for variations in methods of measurement. Initial efforts in the establishment of the PDD have been concentrated on alterations in the acute-phase proteins in individuals with acute and chronic diseases. Even at this early stage in the development of our database, it has proven to be useful as we have found that there appear to be several common acute-phase protein alterations in the plasma and cerebrospinal fluid from patients with Alzheimer's disease, schizophrenia and major depression. Our goal is to provide access to the PDD so that systematic correlations and relationships between disease states can be examined and extended.

Body Fluids↗