PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Discovering empirically conserved amino acid substitution groups in databases of protein families.

This paper introduces a method for identifying empirically conserved amino acid substitution groups. In contrast with existing approaches that view amino acid substitution as a pairwise phenomenon, the method presented here identifies conserved groups of amino acids using a data structure called a conditional distribution matrix. The conditional distribution matrix extends the concept of a pairwise substitution matrix by changing the context of substitution from a single amino acid to a group of amino acids. The matrix tabulates information from a database of protein families that contains numerous aligned positions. Each row in the matrix contains the distribution of amino acids in those aligned positions that contain a given conditioning group of amino acids. The method converts a database of protein families into a conditional distribution matrix and then examines each possible substitution group for evidence of conservation. The algorithm is applied to the BLOCKS and HSSP databases. Twenty amino acid substitution groups are found to be conserved empirically in both databases. These groups provide insight into biochemical properties that are conserved in protein evolution.

Algorithms↗

Database of protein sequence alignments: PIR-ALN.

The Protein Information Resource (PIR) has been maintaining a database of curated protein sequence alignments since 1991. The collection includes superfamily, family and homology domain alignments. CLUSTAL V/W is used to generate multiple sequence alignments and ALNED, an interactive alignment editor, is used to check and correct them. The database has helped in classifying sequences, in defining new homology domains, and in spreading and standardizing protein names, features and keywords among members of a family or superfamily. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. The quarterly and weekly updates can be accessed via the WWW at http://www-nbrf. georgetown.edu/pir/

Databases, Factual↗

SURFACE: a database of protein surface regions for functional annotation.

The SURFACE (SUrface Residues and Functions Annotated, Compared and Evaluated, URL http://cbm.bio.uniroma2.it/surface/) database is a repository of annotated and compared protein surface regions. SURFACE contains the results of a large-scale protein annotation and local structural comparison project. A non-redundant set of protein chains is used to build a database of protein surface patches, defined as putative surface functional sites. Each patch is annotated with sequence and structure-derived information about function or interaction abilities. A new procedure for structure comparison is used to perform an all-versus-all patches comparison. Selection of the results obtained with stringent parameters offers a similarity score that can be used to associate different patches and allows reliable annotation by similarity. Annotation exerted through the comparison of regions of protein surface allows the highlighting of similarities that cannot be recognized by other methods of sequence or structure comparison. A graphic representation of the surface patches, functional annotations and the structural superpositions is available through the web interface.

Algorithms↗

MIPS: a database for protein sequences and complete genomes.

The MIPS group [Munich Information Center for Protein Sequences of the German National Center for Environment and Health (GSF)] at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, is involved in a number of data collection activities, including a comprehensive database of the yeast genome, a database reflecting the progress in sequencing the Arabidopsis thaliana genome, the systematic analysis of other small genomes and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). Through its WWW server (http://www.mips.biochem.mpg.de ) MIPS provides access to a variety of generic databases, including a database of protein families as well as automatically generated data by the systematic application of sequence analysis algorithms. The yeast genome sequence and its related information was also compiled on CD-ROM to provide dynamic interactive access to the 16 chromosomes of the first eukaryotic genome unraveled.

Amino Acid Sequence↗

The Pfam protein families database.

Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the WWW in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgr.ki.se/Pfam/ and in the US at http://pfam.wustl.edu/. The latest version (4.3) of Pfam contains 1815 families. These Pfam families match 63% of proteins in SWISS-PROT 37 and TrEMBL 9. For complete genomes Pfam currently matches up to half of the proteins. Genomic DNA can be directly searched against the Pfam library using the Wise2 package.

Databases, Factual↗

The mammalian protein-protein interaction database and its viewing system that is linked to the main FANTOM2 viewer.

Here, we describe the development of a mammalian protein-protein interaction (PPI) database and of a PPI Viewer application to display protein interaction networks (http://fantom21.gsc.riken.go.jp/PPI/). In the database, we stored the mammalian PPIs identified through our PPI assays (internal PPIs), as well as those we extracted and processed (external PPIs) from publicly available data sources, the DIP and BIND databases and MEDLINE abstracts by using FACTS, a new functional inference and curation system. We integrated the internal and external PPIs into the PPI database, which is linked to the main FANTOM2 viewer. In addition, we incorporated into the PPI Viewer information regarding the luciferase reporter activity of internal PPIs and the data confidence of external PPIs; these data enable visualization and evaluation of the reliability of each interaction. Using the described system, we successfully identified several interactions of biological significance. Therefore, the PPI Viewer is a useful tool for exploring FANTOM2 clone-related protein interactions and their potential effects on signaling and cellular communication.

Animals↗

Integrated approach for manual evaluation of peptides identified by searching protein sequence databases with tandem mass spectra.

Quantitative proteomics relies on accurate protein identification, which often is carried out by automated searching of a sequence database with tandem mass spectra of peptides. When these spectra contain limited information, automated searches may lead to incorrect peptide identifications. It is therefore necessary to validate the identifications by careful manual inspection of the mass spectra. Not only is this task time-consuming, but the reliability of the validation varies with the experience of the analyst. Here, we report a systematic approach to evaluating peptide identifications made by automated search algorithms. The method is based on the principle that the candidate peptide sequence should adequately explain the observed fragment ions. Also, the mass errors of neighboring fragments should be similar. To evaluate our method, we studied tandem mass spectra obtained from tryptic digests of E. coli and HeLa cells. Candidate peptides were identified with the automated search engine Mascot and subjected to the manual validation method. The method found correct peptide identifications that were given low Mascot scores (e.g., 20-25) and incorrect peptide identifications that were given high Mascot scores (e.g., 40-50). The method comprehensively detected false results from searches designed to produce incorrect identifications. Comparison of the tandem mass spectra of synthetic candidate peptides to the spectra obtained from the complex peptide mixtures confirmed the accuracy of the evaluation method. Thus, the evaluation approach described here could help boost the accuracy of protein identification, increase number of peptides identified, and provide a step toward developing a more accurate next-generation algorithm for protein identification.

Algorithms↗

Mass spectrometric sequencing of endotoxin proteins of Bacillus thuringiensis ssp. konkukian extracted from polyacrylamide gels.

The amino acid sequences of the crystal proteins of Bacillus thuringiensis ssp. konkukian strain HL-47 are unknown. We used 1-D denaturing polyacrylamide electrophoresis, nano-ESI-Q-TOF-MS, and protein database searching to analyze these proteins. On SDS-PAGE gels, a preparation of purified crystal proteins exhibited 110, 102, 76, 55, 37, and 30 kDa protein bands. Immunoblotting of the gel with antiserum raised to this preparation revealed that four crystal proteins, of 110, 102, 55, and 37 kDa, reacted with the specific antiserum. The 102-kDa major protein reacted strongly. The other crystal proteins showed weak immunoreactivity. The 102 and 55 kDa proteins were analyzed by ESI-MS. The internal amino acid sequence of the 102-kDa major protein has similarity to the sequences of the surface layer protein of B. thuringiensis ssp. finitimus and B. anthracis. However, the internal amino acid sequences of the 55 kDa protein did not show any homology to proteins in the databases. Proteomic analysis of these proteins leads to the conclusion that the sequence data provided the protein databases of the crystal proteins of the konkukian ssp.

Amino Acid Sequence↗

Scrutineer: a computer program that flexibly seeks and describes motifs and profiles in protein sequence databases.

Scrutineer is an interactive, user-friendly program designed to search for motifs, patterns and profiles in the Swissprot, Protein Identification Resource (PIR) or SeqDb protein sequence databases. Basic capabilities include (i) searches for strings of amino acids with multiple choices at a given position; (ii) searches for strings including variable-length segments and delocalized constraints; (iii) searches over subsets of a database or particular regions within each sequence (e.g. N-terminal one-third); (iv) searches involving secondary structure predictions, physicochemical characteristics, and the like; and (v) searches using aligned sequences as targets with various optional weighting schemes. The various search criteria and hits can be combined and complex targets located. Once the data are loaded into virtual memory, all occurrences in PIR release 22.0 (3.7 x 10(6) amino acids) of a given short string of amino acids (e.g. a hexamer) are found in approximately 36 s. Scrutineer can also describe the entire database, user-specified hits, user-defined regions of sequence and all hits. The source code and accompanying manual are being freely distributed.

Algorithms↗

An update of the DEF database of protein fold class predictions.

An update is given on the Database of Expected Fold classes (DEF) that contains a collection of fold-class predictions made from protein sequences and a mail server that provides new predictions for new sequences. To any given sequence one of 49 fold-classes is chosen to classify the structure related to the sequence with high accuracy. The updated prediction system is developed using data from the new version of the 3D-ALI database of aligned protein structures and thus is giving more reliable and more detailed predictions than the previous DEF system.

Amino Acid Sequence↗

3MOTIF: visualizing conserved protein sequence motifs in the protein structure database.

SUMMARY: 3MOTIF is a web application that visually maps conserved sequence motifs onto three-dimensional protein structures in the Protein Data Bank (PDB; Berman et al., Nucleic Acids Res., 28, 235-242, 2000). Important properties of motifs such as conservation strength and solvent accessible surface area at each position are visually represented on the structure using a variety of color shading schemes. Users can manipulate the displayed motifs using the freely available Chime plugin. AVAILABILITY: http://motif.stanford.edu/3motif/

Amino Acid Motifs↗

Identification and characterization of the peptidic component of the immunomodulatory glycoconjugate Immunoferon.

Inmunoferon is a glycoconjugate of natural origin, formed by the noncovalent association of a protein from Ricinus communis and a polysacharidic moiety, and endowed with immunomodulatory as well as pharmacological activities. This study investigated the nature of polypeptidic component of Inmunoferon. Through biochemical procedures and comparison with protein databases, the isolated protein was identified as the processed form of the seed of Ricinus communis 2S storage polypeptide, which has been termed RicC3. Further analysis of the isolated protein has revealed that it is composed of two different subunits, alpha and beta, which form an heterodimer of high stability and resistance to denaturation, acidic pH and proteolytic cleavage. These findings confirm the excellent properties of the product after oral administration and provide additional support for the pharmacological activities of Inmunoferon.

Adjuvants, Immunologic↗

iGNM: a database of protein functional motions based on Gaussian Network Model.

MOTIVATION: The knowledge of protein structure is not sufficient for understanding and controlling its function. Function is a dynamic property. Although protein structural information has been rapidly accumulating in databases, little effort has been invested to date toward systematically characterizing protein dynamics. The recent success of analytical methods based on elastic network models, and in particular the Gaussian Network Model (GNM), permits us to perform a high-throughput analysis of the collective dynamics of proteins. RESULTS: We computed the GNM dynamics for 20 058 structures from the Protein Data Bank, and generated information on the equilibrium dynamics at the level of individual residues. The results are stored on a web-based system called iGNM and configured so as to permit the users to visualize or download the results through a standard web browser using a simple search engine. Static and animated images for describing the conformational mobility of proteins over a broad range of normal modes are accessible, along with an online calculation engine available for newly deposited structures. A case study of the dynamics of 20 non-homologous hydrolases is presented to illustrate the utility of the iGNM database for identifying key residues that control the cooperative motions and revealing the connection between collective dynamics and catalytic activity.

Algorithms↗

PASS2: an automated database of protein alignments organised as structural superfamilies.

BACKGROUND: The functional selection and three-dimensional structural constraints of proteins in nature often relates to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. Organization of structure-based sequence alignments for distantly related proteins, provides a map of the conserved and critical regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The Protein Alignment organised as Structural Superfamily (PASS2) database represents continuously updated, structural alignments for evolutionary related, sequentially distant proteins. DESCRIPTION: An automated and updated version of PASS2 is, in direct correspondence with SCOP 1.63, consisting of sequences having identity below 40% among themselves. Protein domains have been grouped into 628 multi-member superfamilies and 566 single member superfamilies. Structure-based sequence alignments for the superfamilies have been obtained using COMPARER, while initial equivalencies have been derived from a preliminary superposition using LSQMAN or STAMP 4.0. The final sequence alignments have been annotated for structural features using JOY4.0. The database is supplemented with sequence relatives belonging to different genomes, conserved spatially interacting and structural motifs, probabilistic hidden markov models of superfamilies based on the alignments and useful links to other databases. Probabilistic models and sensitive position specific profiles obtained from reliable superfamily alignments aid annotation of remote homologues and are useful tools in structural and functional genomics. PASS2 presents the phylogeny of its members both based on sequence and structural dissimilarities. Clustering of members allows us to understand diversification of the family members. The search engine has been improved for simpler browsing of the database. CONCLUSIONS: The database resolves alignments among the structural domains consisting of evolutionarily diverged set of sequences. Availability of reliable sequence alignments of distantly related proteins despite poor sequence identity and single-member superfamilies permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. PASS2 is accessible at http://www.ncbs.res.in/~faculty/mini/campass/pass2.html

Amino Acid Sequence↗

Pox proteomics: mass spectrometry analysis and identification of Vaccinia virion proteins.

BACKGROUND: Although many vaccinia virus proteins have been identified and studied in detail, only a few studies have attempted a comprehensive survey of the protein composition of the vaccinia virion. These projects have identified the major proteins of the vaccinia virion, but little has been accomplished to identify the unknown or less abundant proteins. Obtaining a detailed knowledge of the viral proteome of vaccinia virus will be important for advancing our understanding of orthopoxvirus biology, and should facilitate the development of effective antiviral drugs and formulation of vaccines. RESULTS: In order to accomplish this task, purified vaccinia virions were fractionated into a soluble protein enriched fraction (membrane proteins and lateral bodies) and an insoluble protein enriched fraction (virion cores). Each of these fractions was subjected to further fractionation by either sodium dodecyl sulfate-polyacrylamide gel electophoresis, or by reverse phase high performance liquid chromatography. The soluble and insoluble fractions were also analyzed directly with no further separation. The samples were prepared for mass spectrometry analysis by digestion with trypsin. Tryptic digests were analyzed by using either a matrix assisted laser desorption ionization time of flight tandem mass spectrometer, a quadrupole ion trap mass spectrometer, or a quadrupole-time of flight mass spectrometer (the latter two instruments were equipped with electrospray ionization sources). Proteins were identified by searching uninterpreted tandem mass spectra against a vaccinia virus protein database created by our lab and a non-redundant protein database. CONCLUSION: Sixty three vaccinia proteins were identified in the virion particle. The total number of peptides found for each protein ranged from 1 to 62, and the sequence coverage of the proteins ranged from 8.2% to 94.9%. Interestingly, two vaccinia open reading frames were confirmed as being expressed as novel proteins: E6R and L3L.

Amino Acid Sequence↗

ASPD (Artificially Selected Proteins/Peptides Database): a database of proteins and peptides evolved in vitro.

ASPD is a new curated database that incorporates data on full-length proteins, protein domains and peptides that were obtained through in vitro directed evolution processes (mainly by means of phage display). At present, the ASPD database contains data on 195 selection experiments, which were described in 112 original papers. For each experiment, the following information is given: (i) description of the target for binding, (ii) description of the protein or peptide which serves as the template for library construction and description of the native protein which binds the target, (iii) links to the major proteomic databases (SWISS-PROT, PDB, PROSITE and ENZYME), (iv) keywords referring to the biological significance of the experiment, (v) aligned sequences of proteins or peptides retrieved through in vitro evolution and relevant native or constructed sequences, (vi) the number of rounds of selection/amplification and (vii) the number of occurrences of clones with each sequence. The literature data include a full reference, a link to the MEDLINE database and the name of the corresponding author with his email address. ASPD has a user-friendly interface which allows for simple queries using the names of proteins and ligands, as well as keywords describing the biological role of the interaction studied, and also for queries based on authors' names. It is also possible to access the database by means of the SRS system, allowing complex queries. There is a BLAST search tool against the ASPD for looking directly for homologous sequences. Research tools of the ASPD allow the analysis of pairwise correlations in the sequences of proteins and peptides selected against one target. The URL for the ASPD database is http://www.sgi.sscc.ru/mgs/gnw/aspd/.

Animals↗

MinSet: a general approach to derive maximally representative database subsets by using fragment dictionaries and its application to the SCOP database.

MOTIVATION: The size of current protein databases is a challenge for many Bioinformatics applications, both in terms of processing speed and information redundancy. It may be therefore desirable to efficiently reduce the database of interest to a maximally representative subset. RESULTS: The MinSet method employs a combination of a Suffix Tree and a Genetic Algorithm for the generation, selection and assessment of database subsets. The approach is generally applicable to any type of string-encoded data, allowing for a drastic reduction of the database size whilst retaining most of the information contained in the original set. We demonstrate the performance of the method on a database of protein domain structures encoded as strings. We used the SCOP40 domain database by translating protein structures into character strings by means of a structural alphabet and by extracting optimized subsets according to an entropy score that is based on a constant-length fragment dictionary. Therefore, optimized subsets are maximally representative for the distribution and range of local structures. Subsets containing only 10% of the SCOP structure classes show a coverage of >90% for fragments of length 1-4. AVAILABILITY: http://mathbio.nimr.mrc.ac.uk/~jkleinj/MinSet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗