PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include format and content enhancements, cross-references to additional databases, new documentation files and improvements to TrEMBL, a computer-annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDSs) in the EMBL Nucleotide Sequence Database, except the CDSs already included in SWISS-PROT. We also describe the Human Proteomics Initiative (HPI), a major project to annotate all known human sequences according to the quality standards of SWISS-PROT. SWISS-PROT is available at: http://www.expasy.ch/sprot/ and http://www.ebi.ac.uk/swissprot/

Animals↗

Inverse 18O labeling mass spectrometry for the rapid identification of marker/target proteins.

Systematic analysis of proteins is essential in understanding human diseases and their clinical treatments. To achieve the rapid and unambiguous identification of marker or target proteins, a new procedure termed "inverse labeling" is proposed. With this procedure, to evaluate protein expression of a diseased or a drug-treated sample in comparison with a control sample, two converse labeling experiments are performed in parallel. The perturbed sample (by disease or by drug treatment) is labeled in one experiment, whereas the control is labeled in the second experiment. When mixed and analyzed with its unlabeled counterpart for differential comparison using mass spectrometry, a characteristic inverse labeling pattern of mass shift will be observed between the two parallel analyses for proteins that are differentially expressed. In this study, protein labeling is achieved through 18O incorporation into peptides by proteolysis performed in [18O]water. Once the peptides are identified with the characteristic inverse labeling pattern of 18O/16O ion intensity shift, MS data of peptide fingerprints or peptide sequence information can be used to search a protein database for protein identification. The methodology has been applied successfully to two model systems in this study. It permits quick focus on the signals of differentially expressed proteins. It eliminates the detection ambiguities caused by the dynamic range of detection on proteins of extreme changes in expression. It enables the detection of protein modifications responding to perturbation. This strategy can also be extended to other protein-labeling methods, such as chemical or metabolic labeling, to realize the same benefits.

Biomarkers↗

Searching protein structure databases has come of age.

The number of protein structures known in atomic detail has increased from one in 1960 (Kendrew, J.C., Strandberg, B.E., Hart, R.G., Davies, D.R., Phillips, D.C., Shore, V.C. Nature (London) 185:422-427, 1960) to more than 1000 in 1994. The rate at which new structures are being published exceeds one a day as a result of recent advances in protein engineering, crystallography, and spectroscopy. More and more frequently, a newly determined structure is similar in fold to a known one, even when no sequence similarity is detectable. A new generation of computer algorithms has now been developed that allows routine comparison of a protein structure with the database of all known structures. Such structure database searches are already used daily and they are beginning to rival sequence database searches as a tool for discovering biologically interesting relationships.

Algorithms↗

DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions.

The Database of Interacting Proteins (DIP: http://dip.doe-mbi.ucla.edu) is a database that documents experimentally determined protein-protein interactions. It provides the scientific community with an integrated set of tools for browsing and extracting information about protein interaction networks. As of September 2001, the DIP catalogs approximately 11 000 unique interactions among 5900 proteins from >80 organisms; the vast majority from yeast, Helicobacter pylori and human. Tools have been developed that allow users to analyze, visualize and integrate their own experimental data with the information about protein-protein interactions available in the DIP database.

Animals↗

Comparative proteomic analysis of human whole saliva.

Human saliva performs a wide variety of biological functions that are critical for the maintenance of the oral health. Various functions include lubrication, buffering, antimicrobial protection, and the maintenance of mucosal integrity. In addition, whole saliva may be analysed for the diagnosis of human systemic diseases, since it can be readily collected and contains identifiable serum constituents. By using proteomic approach, we have established a reference proteome map of human whole saliva allowing for the resolution of greater than 200 protein spots in a single two-dimensional polyacrylamide gel. Fifty-four protein spots, comprised of 26 different proteins, were identifies using N-terminal sequencing, mass spectrometry, and/or computer matching with protein database. Ten proteins, whose levels were significantly different when bleeding had occurred in the oral cavity, were discussed in this study. These 10 proteins include alpha-1-antrypsin, apolipoprotein A-I, cystatin A, SA, SA-III, and SN, enolase I, hemoglobin beta-chain, thioredoxin peroxiredoxin B, as well as a prolactin-inducible protein. The proteomic approach identifies candidates from human whole saliva that may prove to be of diagnostic and therapeutic significance.

Adult↗

Pfam: a comprehensive database of protein domain families based on seed alignments.

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Amino Acid Sequence↗

TOPOFIT-DB, a database of protein structural alignments based on the TOPOFIT method.

TOPOFIT-DB (T-DB) is a public web-based database of protein structural alignments based on the TOPOFIT method, providing a comprehensive resource for comparative analysis of protein structure families. The TOPOFIT method is based on the discovery of a saturation point on the alignment curve (topomax point) which presents an ability to objectively identify a border between common and variable parts in a protein structural family, providing additional insight into protein comparison and functional annotation. TOPOFIT also effectively detects non-sequential relations between protein structures. T-DB provides users with the convenient ability to retrieve and analyze structural neighbors for a protein; do one-to-all calculation of a user provided structure against the entire current PDB release with T-Server, and pair-wise comparison using the TOPOFIT method through the T-Pair web page. All outputs are reported in various web-based tables and graphics, with automated viewing of the structure-sequence alignments in the Friend software package for complete, detailed analysis. T-DB presents researchers with the opportunity for comprehensive studies of the variability in proteins and is publicly available at http://mozart.bio.neu.edu/topofit/index.php.

Databases, Protein↗

Predicting helical segments in proteins by a helix-coil transition theory with parameters derived from a structural database of proteins.

A novel helix-coil transition theory has been developed. This new theory contains more types of interactions than similar theories developed earlier. The parameters of the models were obtained from a database of 351 nonhomologous proteins. No manual adjustment of the parameters was performed. The interaction parameters obtained in this manner were found to be physically meaningful, consistent with current understanding of helix stabilizing/destabilizing interactions. Novel insights into helix stabilizing/destabilizing interactions have also emerged from this analysis. The theory developed here worked well in sorting out helical residues from amino acid sequences. If the theory was forced to make prediction on every residue of a given amino acid sequence, its performance was the best among ten other secondary structural prediction algorithms in distinguishing helical residues from nonhelical ones. The theory worked even better if one only required it to make prediction on residues that were "predictable" (identifiable by the theory); > 90% predictive reliability could be achieved. The helical residues or segments identified by the helix-coil transition theory can be used as secondary structural contraints to speed up the prediction of the three-dimensional structure of a protein by reducing the dimension of a computational protein folding problem. Possible further improvements of this helix-coil transition theory are also discussed.

Algorithms↗

RSDB: representative protein sequence databases have high information content.

MOTIVATION: Biological sequence databases are highly redundant for two main reasons: 1. various databanks keep redundant sequences with many identical and nearly identical sequences 2. natural sequences often have high sequence identities due to gene duplication. We wanted to know how many sequences can be removed before the databases start losing homology information. Can a database of sequences with mutual sequence identity of 50% or less provide us with the same amount of biological information as the original full database? RESULTS: Comparisons of nine representative sequence databases (RSDB) derived from full protein databanks showed that the information content of sequence databases is not linearly proportional to its size. An RSDB reduced to mutual sequence identity of around 50% (RSDB50) was equivalent to the original full database in terms of the effectiveness of homology searching. It was a third of the full database size which resulted in a six times faster iterative profile searching. The RSDBs are produced at different granularity for efficient homology searching. AVAILABILITY: All the RSDB files generated and the full analysis results are available through internet: ftp://ftp.ebi.ac. uk/pub/contrib/jong/RSDB/http://cyrah.e bi.ac.uk:1111/Proj/Bio/RSDB

Algorithms↗

A database of protein expression in lung cancer.

We have developed a comprehensive approach to identifying molecular changes in lung cancer that includes both genomic and proteomic analyses. The related effort has produced a large amount of data pertaining to gene expression at the RNA and protein levels. As a result, we have constructed a database that contains protein expression data on lung cancer as well as other relevant data including DNA microarray derived data. A large number of proteins that are expressed in different types of lung cancer have been identified and have been correlated with the expression measures for their corresponding genes at the RNA level. The database is intended to facilitate our effort at developing novel classification schemes for lung cancer and the identification of novel markers for early diagnosis.

Biomarkers, Tumor↗

iProClass: an integrated, comprehensive and annotated protein classification database.

The iProClass database is an integrated resource that provides comprehensive family relationships and structural and functional features of proteins, with rich links to various databases. It is extended from ProClass, a protein family database that integrates PIR superfamilies and PROSITE motifs. The iProClass currently consists of more than 200,000 non-redundant PIR and SWISS-PROT proteins organized with more than 28,000 superfamilies, 2600 domains, 1300 motifs, 280 post-translational modification sites and links to more than 30 databases of protein families, structures, functions, genes, genomes, literature and taxonomy. Protein and family summary reports provide rich annotations, including membership information with length, taxonomy and keyword statistics, full family relationships, comprehensive enzyme and PDB cross-references and graphical feature display. The database facilitates classification-driven annotation for protein sequence databases and complete genomes, and supports structural and functional genomic research. The iProClass is implemented in Oracle 8i object-relational system and available for sequence search and report retrieval at http://pir.georgetown.edu/iproclass/.

Databases, Factual↗

Iditis: protein structure database.

The validation, enrichment and organization of the data stored in PDB files is essential for those data to be used accurately and efficiently for modelling, experimental design and the determination of molecular interactions. The Iditis protein structure database has been designed to allow the widest possible range of queries to be performed across all available protein structures. The Iditis database is the most comprehensive protein structure resource currently available, and contains over 500 fields of information describing all publicly deposited protein structures. A custom-written database engine and graphical user interface provide a natural and simple environment for the construction of searches for complex sequence- and structure-based motifs. Extensions and specialized interfaces allow the data generated by the database to used in conjunction with a wide range of applications.

Database Management Systems↗

The CATH protein family database: a resource for structural and functional annotation of genomes.

Over the last decade, there have been huge increases in the numbers of protein sequences and structures determined. In parallel, many methods have been developed for recognising similarities between these proteins, arising from their common evolutionary background, and for clustering such relatives into protein families. Here we review some of the protein family resources available to the biologist and describe how these can be used to provide structural and functional annotations for newly determined sequences. In particular we describe recent developments to the CATH domain database of protein structural families which have facilitated genome annotation and which have also revealed important caveats that must be considered when transferring functional data between homologous proteins.

Databases, Protein↗

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

Human protein reference database as a discovery resource for proteomics.

The rapid pace at which genomic and proteomic data is being generated necessitates the development of tools and resources for managing data that allow integration of information from disparate sources. The Human Protein Reference Database (http://www.hprd.org) is a web-based resource based on open source technologies for protein information about several aspects of human proteins including protein-protein interactions, post-translational modifications, enzyme-substrate relationships and disease associations. This information was derived manually by a critical reading of the published literature by expert biologists and through bioinformatics analyses of the protein sequence. This database will assist in biomedical discoveries by serving as a resource of genomic and proteomic information and providing an integrated view of sequence, structure, function and protein networks in health and disease.

Computational Biology↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence↗

Identification of proteins of the yeast protein map using genetically manipulated strains and peptide-mass fingerprinting.

In this study we used genetically manipulated strains in order to identify polypeptide spots of the protein map of Saccharomyces cerevisiae. Thirty-two novel polypeptide spots were identified using this strategy. They corresponded to the product of 23 different genes. We also explored the possibilities of using peptide-mass fingerprinting for the identification of proteins separated on our gels. According to this strategy, proteins contained in spots are digested with trypsin and the masses of generated peptides are determined by matrix-assisted laser desorption-ionization mass spectrometry (MALDI-MS). The peptide masses are then used to search a yeast protein database for proteins that match the experimental data. Application of this strategy to previously identified polypeptide spots gave evidence of the feasibility of this approach. We also report predictions on the identities of nine unknown spots using MALDI-MS.

Electrophoresis, Gel, Two-Dimensional↗