PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Puzzle pieces defined: locating common packing units in tertiary protein contacts.

Puzzle pieces are defined as small packing units which make up the unique tertiary interactions in proteins. Anti-parallel and perpendicular helix-helix contacts were broken down into basic puzzle-piece pairs in order to study the traits of such contacts: their limited geometry, preferred residue involvement, residue conformation and other common constraints. These traits can then be used for continued comparison of other protein structures, improving models of and designing proteins de novo and, in time, predicting 3D structure from primary sequence. Results from a small (100 proteins) database of anti-parallel helix-helix contacts and from preliminary work on a large database (600 proteins) of perpendicular helix-helix contacts are presented.

Amino Acid Sequence

A contact scoring matrix for qualitative prediction of change in folding of alpha-helices in globular proteins caused by a mutation.

The atomic pairs in contact for atoms from pairs of amino-acid residues on pairs of helices in a protein database consisting of 48 proteins of known tertiary structure from the Brookhaven Protein Data Bank are searched and counted to construct a primary scoring system. Each score in the primary scoring system is weighted further with the possibility of occurrence of each residue pair in the protein database to give a final scoring matrix. Scores for predicting change in folding of alpha-helices in a mutant protein are calculated by assuming that every pair of helices in the protein can closely interact with each other. It is shown that the change in folding of alpha-helices in several mutant proteins are reflected in both the change of the contact scores and the helix geometry calculated.

Databases, Factual

The protein disease database of human body fluids: I. Rationale for the development of this database.

We are developing a relational database to facilitate quantitative and qualitative comparisons of proteins in human body fluids in normal and disease states. For decades researchers and clinicians have been studying proteins in body fluids such as serum, plasma, cerebrospinal fluid and urine. Currently, most clinicians evaluate only a few specific proteins in a body fluid such as plasma when they suspect that a patient has a disease. Now, however, high resolution two-dimensional protein electrophoresis allows the simultaneous evaluation of 1,500 to 3,000 proteins in complex solutions, such as the body fluids. This and other high resolution methods have encouraged us to collect the clinical data for the body fluid proteins into an easily accessed database. For this reason, it has been constructed on the Internet World Wide Web (WWW) under the title Protein Disease Database (PDD). In addition, this database will provide a linkage between the disease-associated protein alterations and images of the appropriate proteins on high-resolution electrophoretic gels of the body fluids. This effort requires the normalization of data to account for variations in methods of measurement. Initial efforts in the establishment of the PDD have been concentrated on alterations in the acute-phase proteins in individuals with acute and chronic diseases. Even at this early stage in the development of our database, it has proven to be useful as we have found that there appear to be several common acute-phase protein alterations in the plasma and cerebrospinal fluid from patients with Alzheimer's disease, schizophrenia and major depression. Our goal is to provide access to the PDD so that systematic correlations and relationships between disease states can be examined and extended.

Body Fluids

Searching protein structure databases has come of age.

The number of protein structures known in atomic detail has increased from one in 1960 (Kendrew, J.C., Strandberg, B.E., Hart, R.G., Davies, D.R., Phillips, D.C., Shore, V.C. Nature (London) 185:422-427, 1960) to more than 1000 in 1994. The rate at which new structures are being published exceeds one a day as a result of recent advances in protein engineering, crystallography, and spectroscopy. More and more frequently, a newly determined structure is similar in fold to a known one, even when no sequence similarity is detectable. A new generation of computer algorithms has now been developed that allows routine comparison of a protein structure with the database of all known structures. Such structure database searches are already used daily and they are beginning to rival sequence database searches as a tool for discovering biologically interesting relationships.

Algorithms

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence

Predicting helical segments in proteins by a helix-coil transition theory with parameters derived from a structural database of proteins.

A novel helix-coil transition theory has been developed. This new theory contains more types of interactions than similar theories developed earlier. The parameters of the models were obtained from a database of 351 nonhomologous proteins. No manual adjustment of the parameters was performed. The interaction parameters obtained in this manner were found to be physically meaningful, consistent with current understanding of helix stabilizing/destabilizing interactions. Novel insights into helix stabilizing/destabilizing interactions have also emerged from this analysis. The theory developed here worked well in sorting out helical residues from amino acid sequences. If the theory was forced to make prediction on every residue of a given amino acid sequence, its performance was the best among ten other secondary structural prediction algorithms in distinguishing helical residues from nonhelical ones. The theory worked even better if one only required it to make prediction on residues that were "predictable" (identifiable by the theory); > 90% predictive reliability could be achieved. The helical residues or segments identified by the helix-coil transition theory can be used as secondary structural contraints to speed up the prediction of the three-dimensional structure of a protein by reducing the dimension of a computational protein folding problem. Possible further improvements of this helix-coil transition theory are also discussed.

Algorithms

Iditis: protein structure database.

The validation, enrichment and organization of the data stored in PDB files is essential for those data to be used accurately and efficiently for modelling, experimental design and the determination of molecular interactions. The Iditis protein structure database has been designed to allow the widest possible range of queries to be performed across all available protein structures. The Iditis database is the most comprehensive protein structure resource currently available, and contains over 500 fields of information describing all publicly deposited protein structures. A custom-written database engine and graphical user interface provide a natural and simple environment for the construction of searches for complex sequence- and structure-based motifs. Extensions and specialized interfaces allow the data generated by the database to used in conjunction with a wide range of applications.

Database Management Systems

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence

Identification of proteins of the yeast protein map using genetically manipulated strains and peptide-mass fingerprinting.

In this study we used genetically manipulated strains in order to identify polypeptide spots of the protein map of Saccharomyces cerevisiae. Thirty-two novel polypeptide spots were identified using this strategy. They corresponded to the product of 23 different genes. We also explored the possibilities of using peptide-mass fingerprinting for the identification of proteins separated on our gels. According to this strategy, proteins contained in spots are digested with trypsin and the masses of generated peptides are determined by matrix-assisted laser desorption-ionization mass spectrometry (MALDI-MS). The peptide masses are then used to search a yeast protein database for proteins that match the experimental data. Application of this strategy to previously identified polypeptide spots gave evidence of the feasibility of this approach. We also report predictions on the identities of nine unknown spots using MALDI-MS.

Electrophoresis, Gel, Two-Dimensional

Identification of macrophage activation associated proteins by two-dimensional gel electrophoresis and microsequencing.

To understand activation in monocytes and macrophages we have studied changes in protein synthesis using the human monocytoid U937 cell line and two-dimensional polyacrylamide gel electrophoresis (2D PAGE) and protein sequencing. U937 cells that had been metabolically labeled during treatment with PMA, LPS, or IFN-gamma showed appreciable increases or decreases in synthesis of 14 proteins when analyzed by 2D PAGE. Although some 20 proteins are reported to be affected by these agents in U937 cells, none of them correspond with the 14 proteins studied here. Of the 14 observed changes, four spots (p41/65, p35/65, p26/44, p20/53) were up-regulated by PMA only, one (p16/44) by LPS only, five spots (p29/47, p26/45, p26/48, p12/47, p10/45) by both LPS and PMA, and, finally, one (p29/45) by all three agents. Two spots (p20/59 and p20/61) were down-regulated by IFN-gamma and one of these spots (p20/59) was up-regulated by LPS. Only one spot (p20/48) was up-regulated by IFN-gamma. Eleven spots with matching mobilities (both M(r) and pI) to those identified in U937 were observed on 2D PAGE gels from human culture derived macrophages. Ten spots from U937 were sequenced by Edman degradation. Two were could not identified from information contained in the available DNA and protein databases and thus represent novel proteins, whereas a further six of the proteins were N-terminally blocked. The remaining two (29/47 and 12/47, respectively) were identified from existing protein databases as translationally controlled tumor protein (TCTP) and cytokeratin. This is the first report of the presence of TCTP in hemopoietic cells and its modulation by PMA or LPS in any cell type. We believe that 2D PAGE and sequencing is a powerful approach for identifying key proteins in macrophage cellular activation.

Amino Acid Sequence

3-D lookup: fast protein structure database searches at 90% reliability.

There are far fewer classes of three-dimensional protein folds than sequence families but the problem of detecting three-dimensional similarities is NP-complete. We present a novel heuristic for identifying 3-D similarities between a query structure and the database of known protein structures. Many methods for structure alignment use a bottom-up approach, identifying first local matches and then solving a combinatorial problem in building up larger clusters of matching substructures. Here, the top-down approach is to start with the global comparison and select a rough superimposition using a fast 3-D lookup of secondary structure motifs. The superimposition is then extended to an alignment of C alpha atoms by an iterative dynamic programming step. An all-against-all comparison of 385 representative proteins (150,000 pair comparisons) took 1 day of computer time on a single R8000 processor. In other words, one query structure is scanned against the database in a matter of minutes. The method is rated at 90% reliability at capturing statistically significant similarities. It is useful as a rapid preprocessor to a comprehensive protein structure database search system.

Amino Acid Sequence

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural three dimensional (3-D) and sequence one dimensional(1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in Swissprot using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 27% of all Swissprot-stored sequences.

Amino Acid Sequence

Discovering empirically conserved amino acid substitution groups in databases of protein families.

This paper introduces a method for identifying empirically conserved amino acid substitution groups. In contrast with existing approaches that view amino acid substitution as a pairwise phenomenon, the method presented here identifies conserved groups of amino acids using a data structure called a conditional distribution matrix. The conditional distribution matrix extends the concept of a pairwise substitution matrix by changing the context of substitution from a single amino acid to a group of amino acids. The matrix tabulates information from a database of protein families that contains numerous aligned positions. Each row in the matrix contains the distribution of amino acids in those aligned positions that contain a given conditioning group of amino acids. The method converts a database of protein families into a conditional distribution matrix and then examines each possible substitution group for evidence of conservation. The algorithm is applied to the BLOCKS and HSSP databases. Twenty amino acid substitution groups are found to be conserved empirically in both databases. These groups provide insight into biochemical properties that are conserved in protein evolution.

Algorithms

Database of protein sequence alignments: PIR-ALN.

The Protein Information Resource (PIR) has been maintaining a database of curated protein sequence alignments since 1991. The collection includes superfamily, family and homology domain alignments. CLUSTAL V/W is used to generate multiple sequence alignments and ALNED, an interactive alignment editor, is used to check and correct them. The database has helped in classifying sequences, in defining new homology domains, and in spreading and standardizing protein names, features and keywords among members of a family or superfamily. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. The quarterly and weekly updates can be accessed via the WWW at http://www-nbrf. georgetown.edu/pir/

Databases, Factual

MIPS: a database for protein sequences and complete genomes.

The MIPS group [Munich Information Center for Protein Sequences of the German National Center for Environment and Health (GSF)] at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, is involved in a number of data collection activities, including a comprehensive database of the yeast genome, a database reflecting the progress in sequencing the Arabidopsis thaliana genome, the systematic analysis of other small genomes and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). Through its WWW server (http://www.mips.biochem.mpg.de ) MIPS provides access to a variety of generic databases, including a database of protein families as well as automatically generated data by the systematic application of sequence analysis algorithms. The yeast genome sequence and its related information was also compiled on CD-ROM to provide dynamic interactive access to the 16 chromosomes of the first eukaryotic genome unraveled.

Amino Acid Sequence

Scrutineer: a computer program that flexibly seeks and describes motifs and profiles in protein sequence databases.

Scrutineer is an interactive, user-friendly program designed to search for motifs, patterns and profiles in the Swissprot, Protein Identification Resource (PIR) or SeqDb protein sequence databases. Basic capabilities include (i) searches for strings of amino acids with multiple choices at a given position; (ii) searches for strings including variable-length segments and delocalized constraints; (iii) searches over subsets of a database or particular regions within each sequence (e.g. N-terminal one-third); (iv) searches involving secondary structure predictions, physicochemical characteristics, and the like; and (v) searches using aligned sequences as targets with various optional weighting schemes. The various search criteria and hits can be combined and complex targets located. Once the data are loaded into virtual memory, all occurrences in PIR release 22.0 (3.7 x 10(6) amino acids) of a given short string of amino acids (e.g. a hexamer) are found in approximately 36 s. Scrutineer can also describe the entire database, user-specified hits, user-defined regions of sequence and all hits. The source code and accompanying manual are being freely distributed.

Algorithms