PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

EMBOPRO--an automatically generated protein sequence database.

For the identification of newly sequenced proteins it is necessary to have a large stock of known proteins for comparison. In this paper we present an automatically generated protein sequence database. The translation program introduced allows a periodical translation of every new release of the EMBL database. Possible errors of the translation are discussed as well as the reliability of the nucleotide sequence data, which turns out to be quite good. A comparison of our translated database with some established ones is given.

Amino Acid Sequence↗

DBAli: a database of protein structure alignments.

SUMMARY: The DBAli database includes approximately 35000 alignments of pairs of protein structures from SCOP (Lo Conte et al., Nucleic Acids Res., 28, 257-259, 2000) and CE (Shindyalov and Bourne, Protein Eng., 11, 739-747, 1998). DBAli is linked to several resources, including Compare3D (Shindyalov and Bourne, http://www.sdsc.edu/pb/software.htm, 1999) and ModView (Ilyin and Sali, http://guitar.rockefeller.edu/ModView/, 2001) for visualizing sequence alignments and structure superpositions. A flexible search of DBAli by protein sequence and structure properties allows construction of subsets of alignments suitable for a number of applications, such as benchmarking of sequence-sequence and sequence-structure alignment methods under a variety of conditions. AVAILABILITY: http://guitar.rockefeller.edu/DBAli/

Computational Biology↗

A protein secondary structure database (PSS).

A protein secondary structure database (PSS) has been designed to correlate the Protein Sequence Database of the PIR-International with the atomic coordinates and bond connectivities database of the Protein Data Bank in the Brookhaven National Laboratory. The present database includes secondary structures determined by X-ray diffraction analysis, but not predicted structures. The database currently contains data from both the Protein Sequence Database and the Protein Data Bank Database, and will encompass the NMR database in the future. The main characteristics of the database are as follows: (1) the secondary structures, sites, regions and domains of structural interest are displayed together with protein primary structures; and (2) the secondary structure of a desired length of peptide fragment is displayed upon request, as are the peptide fragment(s) that correspond to a defined secondary structure. This database also has software to indicate amino acid pairs having hydrogen bonds and to count the occurrence frequency of each pair as well as the conformational parameters widely used in semi-empirical methods of secondary structure prediction.

Amino Acid Sequence↗

A relational database of protein structures designed for flexible enquiries about conformation.

A relational database of protein structure has been developed to enable rapid and flexible enquiries about the occurrence of many aspects of protein architecture. The coordinates of 294 proteins from the Brookhaven Data Bank have been processed by standard computer programs to generate many additional terms that quantify aspects of protein structure. These terms include solvent accessibility, main-chain and side-chain dihedral angles, and secondary structure. In a relational database, the information is stored in tables with columns holding the different terms and rows holding the different entries for the terms. The different relational base tables store the information about the protein coordinate set, the different chains in the protein, the amino acid residues and ligands, the atomic coordinates, the salt bridges, the hydrogen bonds, the disulphide bridges and the close tertiary contacts. The database was established under ORACLE management system. Enquiries are constructed in ORACLE using SQL (structured query language) which is simple to use and alleviates the need for extensive computer programs. A single table can be searched for entries that meet various criteria, e.g. all protein solved to better than a given resolution. The power of the database occurs when several tables, or the entries in a single table, are cross-correlated. For example the dihedral angles of proline in the fourth position in an alpha-helix in high resolution structures can be rapidly obtained. The structural database provides a powerful tool to obtain empirical rules about protein conformation. This database of protein structures is part of a joint project between Birkbeck College and Leeds University to establish an integrated data resource of protein sequences and structures (ISIS) that encodes the complex patterns of residues and coordinates that define protein conformation. The entire data resource (ISIS) will provide a system to guide all areas of protein modelling including structure prediction, site-directed mutagenesis and de novo protein design. The availability of ISIS is described in the paper.

Computer Simulation↗

3Dee: a database of protein structural domains.

UNLABELLED: The 3Dee database is a repository of protein structural domains. It stores alternative domain definitions for the same protein, organises domains into sequence and structural hierarchies, contains non-redundant set(s) of sequences and structures, multiple structure alignments for families of domains, and allows previous versions of the database to be regenerated. AVAILABILITY: 3Dee is accessible on the World Wide Web at the URL http://barton.ebi.ac.uk/servers/3Dee.html.

Databases, Factual↗

MitoProteome: mitochondrial protein sequence database and annotation system.

MitoProteome is an object-relational mitochondrial protein sequence database and annotation system. The initial release contains 847 human mitochondrial protein sequences, derived from public sequence databases and mass spectrometric analysis of highly purified human heart mitochondria. Each sequence is manually annotated with primary function, subfunction and subcellular location, and extensively annotated in an automated process with data extracted from external databases, including gene information from LocusLink and Ensembl; disease information from OMIM; protein-protein interaction data from MINT and DIP; functional domain information from Pfam; protein fingerprints from PRINTS; protein family and family-specific signatures from InterPro; structure data from PDB; mutation data from PMD; BLAST homology data from NCBI NR; and proteins found to be related based on LocusLink and SWISS-PROT references and sequence and taxonomy data. By highly automating the processes of maintaining the MitoProteome Protein List and extracting relevant data from external databases, we are able to present a dynamic database, updated frequently to reflect changes in public resources. The MitoProteome database is publicly available at http://www. mitoproteome.org/. Users may browse and search MitoProteome, and access a complete compilation of data relevant to each protein of interest, cross-linked to external databases.

Computational Biology↗

An object-oriented database for protein structure analysis.

An object-oriented database system has been developed which is being used to store protein structure data. The database can be queried using the logic programming language Prolog or the query language Daplex. Queries retrieve information by navigating through a network of objects which represent the primary, secondary and tertiary structures of proteins. Routines written in both Prolog and Daplex can integrate complex calculations with the retrieval of data from the database, and can also be stored in the database for sharing among users. Thus object-oriented databases are better suited to prototyping applications and answering complex queries about protein structure than relational databases. This system has been used to find loops of varying length and anchor positions when modelling homologous protein structures.

Amino Acid Sequence↗

Development of human protein reference database as an initial platform for approaching systems biology in humans.

Human Protein Reference Database (HPRD) is an object database that integrates a wealth of information relevant to the function of human proteins in health and disease. Data pertaining to thousands of protein-protein interactions, posttranslational modifications, enzyme/substrate relationships, disease associations, tissue expression, and subcellular localization were extracted from the literature for a nonredundant set of 2750 human proteins. Almost all the information was obtained manually by biologists who read and interpreted >300,000 published articles during the annotation process. This database, which has an intuitive query interface allowing easy access to all the features of proteins, was built by using open source technologies and will be freely available at http://www.hprd.org to the academic community. This unified bioinformatics platform will be useful in cataloging and mining the large number of proteomic interactions and alterations that will be discovered in the postgenomic era.

BRCA1 Protein↗

KinG: a database of protein kinases in genomes.

The KinG database is a comprehensive collection of serine/threonine/tyrosine-specific kinases and their homologues identified in various completed genomes using sequence and profile search methods. The database hosted at http://hodgkin. mbu.iisc.ernet.in/ approximately king provides the amino acid sequences, functional domain assignments and classification of gene products containing protein kinase domains. A search tool enabling the retrieval of protein kinases with specified subfamily and domain combinations is one of the key features of the resource. Identification of a kinase catalytic domain in the user's query sequence is possible using another search tool. The occurrence and location of critical catalytic residues if the query has a catalytic kinase domain, recognition of non-kinase domains in the sequence and subfamily classification of the kinase in the query will help in deciphering the biological role of the kinase. This online compilation can also be used to compare the protein kinases of a given subfamily and domain combinations across various genomes. Another exclusive feature of the database is the collection of the Ser/Thr/Tyr protein kinases and similar sequences encoded in the genomes of archaea and bacteria.

Animals↗

The DynDom database of protein domain motions.

UNLABELLED: A relational database has been developed based on the results from the application of the DynDom program to a number of proteins for which multiple X-ray conformers are available. The database is populated via a web-based tool that allows visitors to the website to run the DynDom program server-side by selecting pairs of X-ray conformers by Protein Data Bank code and chain identifier. AVAILABILITY: The website can be found at: http://www.sys.uea.ac.uk/dyndom.

Crystallography↗

Accelerating approximate subsequence search on large protein sequence databases.

Bioinformatics has become an active research area in recent years. The amount of mapped sequences doubles every fourteen months. BLAST has been widely employed for retrieving sequences which has similar portion(s) to a given sequence. However, BLAST has to scan the entire database every time when a query is issued. This can be very time consuming especially when the database is large. In this paper, we study the problem on how to build a persistent index structure for protein sequences to support approximate match. The suffix tree has been proposed as a solution to index sequence database and has been deployed on organizing DNA sequences (Hunt et al. 2001). Unfortunately, it suffers from the problem of "memory bottleneck" that prevents it from being applied efficiently to a large database. The performance even degrades further for protein database due to a larger fanout at each node. Here, we employ an indexing structure, called BASS-tree, to support approximate match in sublinear time on a large protein database. We call this indexing method as sequence approximate match (SAM) index method. The search of approximate matches can be properly directed to the portion in the database with a high potential of matching quickly. It has been demonstrated in our experiments that the potential performance improvement is in an order of magnitude over alternative methods such as the BLAST algorithm and the suffix tree.

Algorithms↗

The PIR-International Protein Sequence Database.

PIR-International is an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. A major objective of PIR-International is to continue the development of the Protein Sequence Database as an essential public resource for protein sequence information. This paper briefly describes the architecture of the Protein Sequence Database and how it and associated data sets are distributed and can be accessed electronically.

Amino Acid Sequence↗

EyeSite: a semi-automated database of protein families in the eye.

The EyeSite is a web-based database of protein families for proteins that function in the eye and their homologous sequences. The resource clusters proteins at different levels of homology in order to facilitate functional annotation of sequences and modelling of proteins from structural homologues. Eye proteins are organized into the tissue types in which they function and are clustered into homologous families using a novel protocol employing the TribeMCL algorithm. Homologous families are further subdivided into sequence clusters for which multiple sequence alignments are generated. Structural annotations from the CATH domain database are provided for nearly 90% of the sequences, and protein family annotations from the Pfam database for approximately 86%. Homology models have also been generated where appropriate. The EyeSite is stored in a relational database and is extensively linked to other online bioinformatics resources to help relate allelic variants, annotations and clinical details to the derived data in the database. The EyeSite is available for online search, sequence information and model retrieval at http://eyesite.cryst.bbk.ac.uk/.

Amino Acid Sequence↗

The ProDom database of protein domain families: more emphasis on 3D.

ProDom is a comprehensive database of protein domain families generated from the global comparison of all available protein sequences. Recent improvements include the use of three-dimensional (3D) information from the SCOP database; a completely redesigned web interface (http://www.toulouse.inra.fr/prodom.html); visualization of ProDom domains on 3D structures; coupling of ProDom analysis with the Geno3D homology modelling server; Bayesian inference of evolutionary scenarios for ProDom families. In addition, we have developed ProDom-SG, a ProDom-based server dedicated to the selection of candidate proteins for structural genomics.

Computer Graphics↗

The SYSTERS Protein Family Database in 2005.

The SYSTERS project aims to provide a meaningful partitioning of the whole protein sequence space by a fully automatic procedure. A refined two-step algorithm assigns each protein to a family and a superfamily. The sequence data underlying SYSTERS release 4 now comprise several protein sequence databases derived from completely sequenced genomes (ENSEMBL, TAIR, SGD and GeneDB), in addition to the comprehensive Swiss-Prot/TrEMBL databases. The SYSTERS web server (http://systers.molgen.mpg.de) provides access to 158 153 SYSTERS protein families. To augment the automatically derived results, information from external databases like Pfam and Gene Ontology are added to the web server. Furthermore, users can retrieve pre-processed analyses of families like multiple alignments and phylogenetic trees. New query options comprise a batch retrieval tool for functional inference about families based on automatic keyword extraction from sequence annotations. A new access point, PhyloMatrix, allows the retrieval of phylogenetic profiles of SYSTERS families across organisms with completely sequenced genomes.

Algorithms↗

Protein Folding Database (PFD 2.0): an online environment for the International Foldeomics Consortium.

The Protein Folding Database (PFD) is a publicly accessible repository of thermodynamic and kinetic protein folding data. Here we describe the first major revision of this work, featuring extensive restructuring that conforms to standards set out by the recently formed International Foldeomics Consortium. The database now adopts standards for data acquisition, analysis and reporting proposed by the consortium, which will facilitate the comparison of folding rates, energies and structure across diverse sets of proteins. Data can now be easily deposited using a rich set of deposition tools. Enhanced search tools allow sophisticated searching and graphical data analysis affords simple data analysis online. PFD can be accessed freely at http://www.foldeomics.org/pfd/.

Databases, Protein↗

Cytoplasmic ribosomal protein genes of the fission yeast Schizosaccharomyces pombe display a unique promoter type: a suggestion for nomenclature of cytoplasmic ribosomal proteins in databases.

We identified 34 new ribosomal protein genes in the Schizosaccharomyces pombe database at the Sanger Centre coding for 30 different ribosomal proteins. All contain the Homol D-box in their promoter. We have shown that Homol D is, in this promoter type, the TATA-analogue. Many promoters contain the Homol E-box, which serves as a proximal activation sequence. Furthermore, comparative sequence analysis revealed a ribosomal protein gene encoding a protein which is the equivalent of the mammalian ribosomal protein L28. The budding yeast Saccharomyces cerevisiae has no L28 equivalent. Over the past 10 years we have isolated and characterized nine ribosomal protein (rp) genes from the fission yeast S.pombe . This endeavor yielded promoters which we have used to investigate the regulation of rp genes. Since eukaryotic ribosomal proteins are remarkably conserved and several rp genes of the budding yeast S.cerevisiae were sequenced in 1985, we probed DNA fragments encoding S.cerevisiae ribosomal proteins with genomic libraries of S.pombe . The deduced amino acid sequence of the different isolated rp genes of fission yeast share between 65 and 85% identical amino acids with their counterparts of budding yeast.

Amino Acid Sequence↗

Progress with the PRINTS protein fingerprint database.

PRINTS is a compendium of protein motif 'fingerprints' derived from the OWL composite sequence database. Fingerprints are groups of motifs within sequence alignments whose conserved nature allows them to be used as signatures of family membership. To date, 400 fingerprints have been constructed and stored in Prints, the size of which has doubled in the last year. The current version, 9.0, encodes approximately 2000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. Fingerprints inherently offer improved diagnostic reliability over single motif methods by virtue of the mutual context provided by motif neighbours. PRINTS thus provides a useful adjunct to the widely used PROSITE dictionary of patterns. The database is now accessible via the Database Browser on the UCL Bioinformatics server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser .

Amino Acid Sequence↗