PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Searching protein structure databases has come of age.

The number of protein structures known in atomic detail has increased from one in 1960 (Kendrew, J.C., Strandberg, B.E., Hart, R.G., Davies, D.R., Phillips, D.C., Shore, V.C. Nature (London) 185:422-427, 1960) to more than 1000 in 1994. The rate at which new structures are being published exceeds one a day as a result of recent advances in protein engineering, crystallography, and spectroscopy. More and more frequently, a newly determined structure is similar in fold to a known one, even when no sequence similarity is detectable. A new generation of computer algorithms has now been developed that allows routine comparison of a protein structure with the database of all known structures. Such structure database searches are already used daily and they are beginning to rival sequence database searches as a tool for discovering biologically interesting relationships.

Algorithms

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence

Predicting helical segments in proteins by a helix-coil transition theory with parameters derived from a structural database of proteins.

A novel helix-coil transition theory has been developed. This new theory contains more types of interactions than similar theories developed earlier. The parameters of the models were obtained from a database of 351 nonhomologous proteins. No manual adjustment of the parameters was performed. The interaction parameters obtained in this manner were found to be physically meaningful, consistent with current understanding of helix stabilizing/destabilizing interactions. Novel insights into helix stabilizing/destabilizing interactions have also emerged from this analysis. The theory developed here worked well in sorting out helical residues from amino acid sequences. If the theory was forced to make prediction on every residue of a given amino acid sequence, its performance was the best among ten other secondary structural prediction algorithms in distinguishing helical residues from nonhelical ones. The theory worked even better if one only required it to make prediction on residues that were "predictable" (identifiable by the theory); > 90% predictive reliability could be achieved. The helical residues or segments identified by the helix-coil transition theory can be used as secondary structural contraints to speed up the prediction of the three-dimensional structure of a protein by reducing the dimension of a computational protein folding problem. Possible further improvements of this helix-coil transition theory are also discussed.

Algorithms

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence

Identification of proteins of the yeast protein map using genetically manipulated strains and peptide-mass fingerprinting.

In this study we used genetically manipulated strains in order to identify polypeptide spots of the protein map of Saccharomyces cerevisiae. Thirty-two novel polypeptide spots were identified using this strategy. They corresponded to the product of 23 different genes. We also explored the possibilities of using peptide-mass fingerprinting for the identification of proteins separated on our gels. According to this strategy, proteins contained in spots are digested with trypsin and the masses of generated peptides are determined by matrix-assisted laser desorption-ionization mass spectrometry (MALDI-MS). The peptide masses are then used to search a yeast protein database for proteins that match the experimental data. Application of this strategy to previously identified polypeptide spots gave evidence of the feasibility of this approach. We also report predictions on the identities of nine unknown spots using MALDI-MS.

Electrophoresis, Gel, Two-Dimensional

Identification of macrophage activation associated proteins by two-dimensional gel electrophoresis and microsequencing.

To understand activation in monocytes and macrophages we have studied changes in protein synthesis using the human monocytoid U937 cell line and two-dimensional polyacrylamide gel electrophoresis (2D PAGE) and protein sequencing. U937 cells that had been metabolically labeled during treatment with PMA, LPS, or IFN-gamma showed appreciable increases or decreases in synthesis of 14 proteins when analyzed by 2D PAGE. Although some 20 proteins are reported to be affected by these agents in U937 cells, none of them correspond with the 14 proteins studied here. Of the 14 observed changes, four spots (p41/65, p35/65, p26/44, p20/53) were up-regulated by PMA only, one (p16/44) by LPS only, five spots (p29/47, p26/45, p26/48, p12/47, p10/45) by both LPS and PMA, and, finally, one (p29/45) by all three agents. Two spots (p20/59 and p20/61) were down-regulated by IFN-gamma and one of these spots (p20/59) was up-regulated by LPS. Only one spot (p20/48) was up-regulated by IFN-gamma. Eleven spots with matching mobilities (both M(r) and pI) to those identified in U937 were observed on 2D PAGE gels from human culture derived macrophages. Ten spots from U937 were sequenced by Edman degradation. Two were could not identified from information contained in the available DNA and protein databases and thus represent novel proteins, whereas a further six of the proteins were N-terminally blocked. The remaining two (29/47 and 12/47, respectively) were identified from existing protein databases as translationally controlled tumor protein (TCTP) and cytokeratin. This is the first report of the presence of TCTP in hemopoietic cells and its modulation by PMA or LPS in any cell type. We believe that 2D PAGE and sequencing is a powerful approach for identifying key proteins in macrophage cellular activation.

Amino Acid Sequence

3-D lookup: fast protein structure database searches at 90% reliability.

There are far fewer classes of three-dimensional protein folds than sequence families but the problem of detecting three-dimensional similarities is NP-complete. We present a novel heuristic for identifying 3-D similarities between a query structure and the database of known protein structures. Many methods for structure alignment use a bottom-up approach, identifying first local matches and then solving a combinatorial problem in building up larger clusters of matching substructures. Here, the top-down approach is to start with the global comparison and select a rough superimposition using a fast 3-D lookup of secondary structure motifs. The superimposition is then extended to an alignment of C alpha atoms by an iterative dynamic programming step. An all-against-all comparison of 385 representative proteins (150,000 pair comparisons) took 1 day of computer time on a single R8000 processor. In other words, one query structure is scanned against the database in a matter of minutes. The method is rated at 90% reliability at capturing statistically significant similarities. It is useful as a rapid preprocessor to a comprehensive protein structure database search system.

Amino Acid Sequence

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural three dimensional (3-D) and sequence one dimensional(1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in Swissprot using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 27% of all Swissprot-stored sequences.

Amino Acid Sequence

Discovering empirically conserved amino acid substitution groups in databases of protein families.

This paper introduces a method for identifying empirically conserved amino acid substitution groups. In contrast with existing approaches that view amino acid substitution as a pairwise phenomenon, the method presented here identifies conserved groups of amino acids using a data structure called a conditional distribution matrix. The conditional distribution matrix extends the concept of a pairwise substitution matrix by changing the context of substitution from a single amino acid to a group of amino acids. The matrix tabulates information from a database of protein families that contains numerous aligned positions. Each row in the matrix contains the distribution of amino acids in those aligned positions that contain a given conditioning group of amino acids. The method converts a database of protein families into a conditional distribution matrix and then examines each possible substitution group for evidence of conservation. The algorithm is applied to the BLOCKS and HSSP databases. Twenty amino acid substitution groups are found to be conserved empirically in both databases. These groups provide insight into biochemical properties that are conserved in protein evolution.

Algorithms

MIPS: a database for protein sequences and complete genomes.

The MIPS group [Munich Information Center for Protein Sequences of the German National Center for Environment and Health (GSF)] at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, is involved in a number of data collection activities, including a comprehensive database of the yeast genome, a database reflecting the progress in sequencing the Arabidopsis thaliana genome, the systematic analysis of other small genomes and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). Through its WWW server (http://www.mips.biochem.mpg.de ) MIPS provides access to a variety of generic databases, including a database of protein families as well as automatically generated data by the systematic application of sequence analysis algorithms. The yeast genome sequence and its related information was also compiled on CD-ROM to provide dynamic interactive access to the 16 chromosomes of the first eukaryotic genome unraveled.

Amino Acid Sequence

Scrutineer: a computer program that flexibly seeks and describes motifs and profiles in protein sequence databases.

Scrutineer is an interactive, user-friendly program designed to search for motifs, patterns and profiles in the Swissprot, Protein Identification Resource (PIR) or SeqDb protein sequence databases. Basic capabilities include (i) searches for strings of amino acids with multiple choices at a given position; (ii) searches for strings including variable-length segments and delocalized constraints; (iii) searches over subsets of a database or particular regions within each sequence (e.g. N-terminal one-third); (iv) searches involving secondary structure predictions, physicochemical characteristics, and the like; and (v) searches using aligned sequences as targets with various optional weighting schemes. The various search criteria and hits can be combined and complex targets located. Once the data are loaded into virtual memory, all occurrences in PIR release 22.0 (3.7 x 10(6) amino acids) of a given short string of amino acids (e.g. a hexamer) are found in approximately 36 s. Scrutineer can also describe the entire database, user-specified hits, user-defined regions of sequence and all hits. The source code and accompanying manual are being freely distributed.

Algorithms

An update of the DEF database of protein fold class predictions.

An update is given on the Database of Expected Fold classes (DEF) that contains a collection of fold-class predictions made from protein sequences and a mail server that provides new predictions for new sequences. To any given sequence one of 49 fold-classes is chosen to classify the structure related to the sequence with high accuracy. The updated prediction system is developed using data from the new version of the 3D-ALI database of aligned protein structures and thus is giving more reliable and more detailed predictions than the previous DEF system.

Amino Acid Sequence

Molecular cloning and expression of a novel keratinocyte protein (psoriasis-associated fatty acid-binding protein [PA-FABP]) that is highly up-regulated in psoriatic skin and that shares similarity to fatty acid-binding proteins.

Analysis by means of two-dimensional (2D) gel electrophoresis of the protein patterns of normal and psoriatic unfractionated non-cultured keratinocytes has revealed a few low-molecular-weight proteins that are highly up-regulated in psoriatic skin. These include psoriasin; calgranulin B, also known as MRP 14, L1, or calprotectin; calgranulin A or MRP 8; and cystatin A or stefin A. Here, we have cloned and sequenced the cDNA (clone 1592) encoding a new member of this group of low-molecular-weight proteins [isoelectric focusing (IEF) SSP 3007 in the keratinocyte 2D gel protein database] that we have termed PA-FABP (psoriasis-associated fatty acid-binding protein). The deduced sequence predicted a protein with molecular weight of 15,164 daltons and a calculated pI of 6.96, values that are close to those recorded in the keratinocyte 2D gel protein database. The protein comigrated with PA-FABP as determined by 2D gel analysis of [35S]-methionine-labeled proteins expressed by transformed human amnion (AMA) cells transfected with clone 1592 using the vaccinia virus expression system and reacted with a rabbit polyclonal antibody raised against 2D gel purified PA-FABP. Structural analysis of the amino acid sequence revealed 48%, 52%, and 56% identity to known low-molecular-weight fatty acid-binding proteins belonging to the FABP family. Northern blot analysis showed that PA-FABP mRNA is indeed highly up-regulated in psoriatic keratinocytes. The transcript is present in human cell lines of epithelial and lymphoid (Molt 4) origin but cannot be detected in normal or SV40 transformed MRC-5 fibroblasts. 2D gel protein analysis of normal primary keratinocytes cultured for at least 8 d under conditions that promoted incomplete terminal differentiation [serum-free keratinocyte (SFK) medium supplemented with epidermal growth factor (EGF), pituitary extract, and 10% fetal calf serum] revealed a strong up-regulation of PA-FABP, psoriasin, calgranulins A and B, and a few other proteins that are highly expressed in psoriatic skin. The levels of these proteins exceeded by far those observed in non-cultured normal keratinocytes implying that the cultured cells have followed an altered pattern of differentiation that resembles--at least in part--that of non-cultured psoriatic keratinocytes. The implications of these results for the study of psoriasis are discussed.

Amino Acid Sequence

Construction of HSC-2DPAGE: a two-dimensional gel electrophoresis database of heart proteins.

The dissemination of information relating to the characterisation of proteins from two-dimensional electrophoresis (2-DE) gel databases is essential for their effective utilisation in the study of protein expression in cell biology. Since the inception of the World Wide Web and the pioneering development of SWISS-2DPAGE as a tool for retrieving information on proteins separated by 2-DE, the Internet has become the method of choice for disseminating and accessing information on 2-DE protein databases. At Harefield we have established HSC-2DPAGE which is an advanced interface for accessing protein database relating to heart disease. The Web site currently includes databases of proteins from human, dog and rat ventricular tissue and a human endothelial cell line. The databases are searchable individually or as a whole by remote keyword searches. Each database is represented by both synthetic (computer generated) and real (scanned gel) clickable images upon which characterised protein spots are highlighted by hyperlinked symbols. The database conforms to all the rules proposed for federated 2-DE protein databases and individual protein entries are linked to other protein databases such as SWISS-PROT by active cross-references. This paper describes the construction of HSC-2DPAGE, its maintenance, and access via the Internet.

Animals

The Protein Information Resource (PIR) and the PIR-International Protein Sequence Database.

From its origin, the PIR has aspired to support research in computational biology and genomics through the compilation of a comprehensive, quality controlled and well-organized protein sequence information resource. The resource originated with the pioneering work of the late Margaret O. Dayhoff in the early 1960s. Since 1988, the Protein Sequence Database has been maintained collaboratively by PIR-International, an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. The work of the resource is widely distributed and is available on the World Wide Web, via FTP, E-mail server, CD-ROM and magnetic media. It is widely redistributed and incorporated into many other protein sequence data compilations including SWISS-PROT and theEntrezsystem of the NCBI.

Amino Acid Sequence

Secondary structure-based profiles: use of structure-conserving scoring tables in searching protein sequence databases for structural similarities.

The profile method, for detecting distantly related proteins by sequence comparison, has been extended to incorporate secondary structure information from known X-ray structures. The sequence of a known structure is aligned to sequences of other members of a given folding class. From the known structure, the secondary structure (alpha-helix, beta-strand or "other") is assigned to each position of the aligned sequences. As in the standard profile method, a position-dependent scoring table, termed a profile, is calculated from the aligned sequences. However, rather than using the standard Dayhoff mutation table in calculating the profile, we use distinct amino acid mutation tables for residues in alpha-helices, beta-strands or other secondary structures to calculate the profile. In addition, we also distinguish between internal and external residues. With this new secondary structure-based profile method, we created a profile for eight-stranded, antiparallel beta barrels of the insecticyanin folding class. It is based on the sequences of retinol-binding protein, insecticyanin and beta-lactoglobulin. Scanning the sequence database with this profile, it was possible to detect the sequence of avidin. The structure of streptavidin is known, and it appears to be distantly related to the antiparallel beta barrels. Also detected is the sequence of complement component C8, which we therefore predict to be a member of this folding class.

Amino Acid Sequence