PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Protein identification from two-dimensional gel electrophoresis analysis of Klebsiella pneumoniae by combined use of mass spectrometry data and raw genome sequences.

Separation of proteins by two-dimensional gel electrophoresis (2-DE) coupled with identification of proteins through peptide mass fingerprinting (PMF) by matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS) is the widely used technique for proteomic analysis. This approach relies, however, on the presence of the proteins studied in public-accessible protein databases or the availability of annotated genome sequences of an organism. In this work, we investigated the reliability of using raw genome sequences for identifying proteins by PMF without the need of additional information such as amino acid sequences. The method is demonstrated for proteomic analysis of Klebsiella pneumoniae grown anaerobically on glycerol. For 197 spots excised from 2-DE gels and submitted for mass spectrometric analysis 164 spots were clearly identified as 122 individual proteins. 95% of the 164 spots can be successfully identified merely by using peptide mass fingerprints and a strain-specific protein database (ProtKpn) constructed from the raw genome sequences of K. pneumoniae. Cross-species protein searching in the public databases mainly resulted in the identification of 57% of the 66 high expressed protein spots in comparison to 97% by using the ProtKpn database. 10 dha regulon related proteins that are essential for the initial enzymatic steps of anaerobic glycerol metabolism were successfully identified using the ProtKpn database, whereas none of them could be identified by cross-species searching. In conclusion, the use of strain-specific protein database constructed from raw genome sequences makes it possible to reliably identify most of the proteins from 2-DE analysis simply through peptide mass fingerprinting.

Journal Article↗

The PANTHER database of protein families, subfamilies, functions and pathways.

PANTHER is a large collection of protein families that have been subdivided into functionally related subfamilies, using human expertise. These subfamilies model the divergence of specific functions within protein families, allowing more accurate association with function (ontology terms and pathways), as well as inference of amino acids important for functional specificity. Hidden Markov models (HMMs) are built for each family and subfamily for classifying additional protein sequences. The latest version, 5.0, contains 6683 protein families, divided into 31,705 subfamilies, covering approximately 90% of mammalian protein-coding genes. PANTHER 5.0 includes a number of significant improvements over previous versions, most notably (i) representation of pathways (primarily signaling pathways) and association with subfamilies and individual protein sequences; (ii) an improved methodology for defining the PANTHER families and subfamilies, and for building the HMMs; (iii) resources for scoring sequences against PANTHER HMMs both over the web and locally; and (iv) a number of new web resources to facilitate analysis of large gene lists, including data generated from high-throughput expression experiments. Efforts are underway to add PANTHER to the InterPro suite of databases, and to make PANTHER consistent with the PIRSF database. PANTHER is now publicly available without restriction at http://panther.appliedbiosystems.com.

Animals↗

Automated protein sequence database classification. II. Delineation Of domain boundaries from sequence similarities.

MOTIVATION: Decomposing each protein into modular domains is a basic prerequisite to classify accurately structural units in biological molecules. Boundaries between domains are indicated by two similar amino acid sequence segments located within the same protein (repeats) or within homologous proteins at notably different distances from their respective N- or C-termini. RESULTS: We have developed an automated method that combines such positional constraints derived from various detected pairwise sequence similarities to delineate the modular organization of proteins. The procedure has been applied to a non-redundant data set of 26 990 proteins whose sequences were taken from the PIR and SWISS-PROT databanks and shared <60% sequence identity amongst pairs. The resultant clustering, delineation and multiple alignment of 24 380 sequence fragments yielded a new database of 4364 domain families. Comparison of the domain collection with that of PRODOM indicates a clear improvement in the number and size of domain families, domain boundaries and multiple sequence alignments. The accuracy and sensitivity of the method are illustrated by results obtained for ankyrin-like repeats and EGF-like modules. AVAILABILITY: The resulting database, called DOMO, is available through the database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr

Algorithms↗

Human genome protein function database.

A database which focuses on the normal functions of the currently-known protein products of the Human Genome was constructed. Information is stored as text, figures, tables, and diagrams. The program contains built-in functions to modify, update, categorize, hypertext, search, create reports, and establish links to other databases. The semi-automated categorization feature of the database program was used to classify these proteins in terms of biomedical functions.

Databases, Factual↗

Characterization of ribosomal proteins as biomarkers for matrix-assisted laser desorption/ionization mass spectral identification of Lactobacillus plantarum.

For rapid identification of bacteria by matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS), a bioinformatics approach using ribosomal subunit proteins as biomarkers has been proposed. This method compares the observed masses for biomarkers with calculated masses as predicted from the amino acid sequences registered on protein databases. To evaluate this approach, the expressed ribosomal proteins of a genome-sequenced bacterium, Lactobacillus plantarum NCIMB 8826, were characterized as a model sample. The protein expression of 42 ribosomal subunit proteins, together with 10 ribosome-associated proteins in the isolated ribosome fraction, was confirmed through two-dimensional gel electrophoresis combined with peptide mass fingerprinting. The observed masses of the proteins in the isolated ribosome fraction were then determined by MALDI-MS. We preliminarily selected 44 biomarkers whose observed masses were matched with the calculated masses predicted from the amino acid sequence registered in the protein databases by considering N-terminal methionine loss only. Of these, the finally selected reliable biomarkers were 34 proteins including 31 ribosomal subunit proteins and 3 ribosome-associated proteins that could be observed in the MALDI mass spectra of the cell lysate sample. These biomarkers were usable in MALDI-MS characterization of two industrial L. plantarum cultures.

Bacterial Proteins↗

EST2Prot: mapping EST sequences to proteins.

BACKGROUND: EST libraries are used in various biological studies, from microarray experiments to proteomic and genetic screens. These libraries usually contain many uncharacterized ESTs that are typically ignored since they cannot be mapped to known genes. Consequently, new discoveries are possibly overlooked. RESULTS: We describe a system (EST2Prot) that uses multiple elements to map EST sequences to their corresponding protein products. EST2Prot uses UniGene clusters, substring analysis, information about protein coding regions in existing DNA sequences and protein database searches to detect protein products related to a query EST sequence. Gene Ontology terms, Swiss-Prot keywords, and protein similarity data are used to map the ESTs to functional descriptors. CONCLUSION: EST2Prot extends and significantly enriches the popular UniGene mapping by utilizing multiple relations between known biological entities. It produces a mapping between ESTs and proteins in real-time through a simple web-interface. The system is part of the Biozon database and is accessible at http://biozon.org/tools/est/.

Animals↗

Expectations from structural genomics revisited: an analysis of structural genomics targets.

BACKGROUND: Current structural genomics projects are being driven by two main goals; to produce a representative set of protein folds that could be used as templates for comparative modeling purposes, and to provide insight into the function of the currently unannotated protein sequences. Such projects may reveal that a newly determined protein structure shares structural similarity with a previously observed structure or that it is a novel fold. The manner in which structure can be used to suggest the function of a protein will depend on the number and diversity of homologous sequences and the extent to which these sequences are functionally characterized. METHOD AND RESULTS: Using sequence searching methods, we analyzed structural genomics target sequences to ascertain if they were members of functionally characterized protein families, protein families of unknown function, or orphan sequences. This analysis provided an indication of what could be expected to emerge from structural genomics projects. Matches were found to approximately 25% of the current functionally unannotated protein families in the PFAM database (protein families database of alignments and hidden Markov models). The 16% of strict orphan sequences will be the most problematic if their structures reveal novel folds. However, out of the remaining target sequences that match families whose members are largely of unknown function, 28% are particularly interesting in that they are part of protein families with considerable sequence diversity. CONCLUSION: The determination of a new structure of a member of these families is likely to offer considerable insight into possible functional roles of these proteins even if it is a new fold. Mapping the sequence conservation onto the structure may reveal functionally important residues for further study by experimental methods.

Databases, Protein↗

Efficient similarity search in protein structure databases by k-clique hashing.

MOTIVATION: Graph-based clique-detection techniques are widely used for the recognition of common substructures in proteins. They permit the detection of resemblances that are independent of sequence or fold homologies and are also able to handle conformational flexibility. Their high computational complexity is often a limiting factor and prevents a detailed and fine-grained modeling of the protein structure. RESULTS: We present an efficient two-step method that significantly speeds up the detection of common substructures, especially when used to screen larger databases. It combines the advantages from both clique-detection and geometric hashing. The method is applied to an established approach for the comparison of protein binding-pockets, and some empirical results are presented. AVAILABILITY: Upon request from the authors.

Algorithms↗

The TIGRFAMs database of protein families.

TIGRFAMs is a collection of manually curated protein families consisting of hidden Markov models (HMMs), multiple sequence alignments, commentary, Gene Ontology (GO) assignments, literature references and pointers to related TIGRFAMs, Pfam and InterPro models. These models are designed to support both automated and manually curated annotation of genomes. TIGRFAMs contains models of full-length proteins and shorter regions at the levels of superfamilies, subfamilies and equivalogs, where equivalogs are sets of homologous proteins conserved with respect to function since their last common ancestor. The scope of each model is set by raising or lowering cutoff scores and choosing members of the seed alignment to group proteins sharing specific function (equivalog) or more general properties. The overall goal is to provide information with maximum utility for the annotation process. TIGRFAMs is thus complementary to Pfam, whose models typically achieve broad coverage across distant homologs but end at the boundaries of conserved structural domains. The database currently contains over 1600 protein families. TIGRFAMs is available for searching or downloading at www.tigr.org/TIGRFAMs.

Animals↗

Effects of formaldehyde inhalation on lung of rats.

OBJECTIVE: To analyze protein changes in the lung of Wistar rats exposed to gaseous formaldehyde (FA) at 32-37 mg/m3 for 4 h/day for 15 days using proteomics technique. METHODS: Lung samples were solubilized and separated by two-dimensional electrophoresis (2-DE), and gel patterns were scanned and analyzed for detection of differently expressed protein spots. These protein spots were identified by MALDI-TOF-MS and NCBInr protein database searching. RESULTS: Four proteins were altered significantly in 32-37 mg/m3 FA group, with 3 proteins up-regulated, 1 protein down-regulated. The 4 proteins were identified as aldose reductase, LIM protein, glyceraldehyde-3-phosphate dehydrogenase, and chloride intracellular channel 3. CONCLUSION: The four proteins are related to cell proliferation induced by FA and defense reaction of anti-oxidation. Proteomics is a powerful tool in research of environmental health, and has prospects in search for protein markers for disease diagnosis and monitoring.

Administration, Inhalation↗

Comparison of vacuum matrix-assisted laser desorption/ionization (MALDI) and atmospheric pressure MALDI (AP-MALDI) tandem mass spectrometry of 2-dimensional separated and trypsin-digested glomerular proteins for database search derived identification.

Mass spectrometric based sequencing of enzymatic generated peptides is widely used to obtain specific sequence tags allowing the unambiguous identification of proteins. In the present study, two types of desorption/ionization techniques combined with different modes of ion dissociation, namely vacuum matrix-assisted laser desorption/ionization (vMALDI) high energy collision induced dissociation (CID) and post-source decay (PSD) as well as atmospheric pressure (AP)-MALDI low energy CID, were applied for the fragmentation of singly protonated peptide ions, which were derived from two-dimensional separated, silver-stained and trypsin-digested hydrophilic as well as hydrophobic glomerular proteins. Thereby, defined properties of the individual fragmentation pattern generated by the specified modes could be observed. Furthermore, the compatibility of the varying PSD and CID (MS/MS) data with database search derived identification using two public accessible search algorithms has been evaluated. The peptide sequence tag information obtained by PSD and high energy CID enabled in the majority of cases an unambiguous identification. In contrast, part of the data obtained by low energy CID were not assignable using similar search parameters and therefore no clear results were obtainable. The knowledge of the properties of available MALDI-based fragmentation techniques presents an important factor for data interpretation using public accessible search algorithms and moreover for the identification of two-dimensional gel separated proteins.

Algorithms↗

Potential for false positive identifications from large databases through tandem mass spectrometry.

The biomedical research community at large is increasingly employing shotgun proteomics for large-scale identification of proteins from enzymatic digests. Typically, the approach used to identify proteins and peptides from tandem mass spectral data is based on the matching of experimentally generated tandem mass spectra to the theoretical best match from a protein database. Here, we present the potential difficulties of using such an approach without statistical consideration of the false positive rate, especially when large databases, as are encountered in eukaryotes are considered. This is illustrated by searching a dataset generated from a multidimensional separation of a eukaryotic tryptic digest against an in silico generated random protein database, which generated a significant number of positive matches, even when previously suggested score filtering criteria are used.

Algorithms↗

A mass spectrometry-based proteomic approach to study Marek's Disease Virus gene expression.

Marek's Disease Virus (MDV) is an avian herpesvirus that causes a lymphoproliferative disorder in chickens. MDV transitions between a lytic phase in which new viruses are produced and a latent phase in which the virus lays dormant. The mechanism controlling this lytic-to-latent switch remains unclear. To better understand the lytic phase of MDV infection, a mass spectrometry-based strategy was developed to identify viral proteins and to qualitatively examine their abundance in lytically infected chicken embryo fibroblast (CEF) cells. A combination of strong cation exchange chromatography (SCXC) and microcapillary reversed-phase liquid chromatography-tandem mass spectrometry (murpLC/MS/MS) was used to resolve peptides from tryptic digests of MDV-infected CEF cell lysates. Peptides were identified by searching the tandem mass spectra against a protein database containing both MDV proteins and all currently available Gallus gallus proteins using the SEQUEST algorithm. A total of 427 MDV peptides, corresponding to 82 unique proteins, were identified, with 56 of them detected with at least two unique peptides. Overall, nearly 80% of all putative MDV proteins expressed in infected CEF cells were identified. We anticipate that this approach will be a viable method for determining how viral and host proteome changes occurring in Marek's Disease pathogenesis regulate the switch between the lytic and latent phases of the MDV life cycle.

Animals↗

Characterization of the protein subset desorbed by MALDI from whole bacterial cells.

This study characterizes various features of the proteins that are detected in MALDI mass spectra when whole bacteria cells are analyzed, in an effort to understand why some proteins are successfully detected and many others are not. Forty peaks observed in the mass range 4,000-20,000 Da in the spectra of Escherichia coli K-12 and 11775 are tentatively assigned to proteins in a protein database, and these proteins are characterized by cell location, copy number, pI, and hydropathicity. Those detected originate in the cytosol and generally share the traits of high abundance within the cell, strong bacisity, and medium hydrophilicity.

Bacterial Proteins↗

Comparative genomics of the pennate diatom Phaeodactylum tricornutum.

Diatoms are one of the most important constituents of phytoplankton communities in aquatic environments, but in spite of this, only recently have large-scale diatom-sequencing projects been undertaken. With the genome of the centric species Thalassiosira pseudonana available since mid-2004, accumulating sequence information for a pennate model species appears a natural subsequent aim. We have generated over 12,000 expressed sequence tags (ESTs) from the pennate diatom Phaeodactylum tricornutum, and upon assembly into a nonredundant set, 5,108 sequences were obtained. Significant similarity (E < 1E-04) to entries in the GenBank nonredundant protein database, the COG profile database, and the Pfam protein domains database were detected, respectively, in 45.0%, 21.5%, and 37.1% of the nonredundant collection of sequences. This information was employed to functionally annotate the P. tricornutum nonredundant set and to create an internet-accessible queryable diatom EST database. The nonredundant collection was then compared to the putative complete proteomes of the green alga Chlamydomonas reinhardtii, the red alga Cyanidioschyzon merolae, and the centric diatom T. pseudonana. A number of intriguing differences were identified between the pennate and the centric diatoms concerning activities of relevance for general cell metabolism, e.g. genes involved in carbon-concentrating mechanisms, cytosolic acetyl-Coenzyme A production, and fructose-1,6-bisphosphate metabolism. Finally, codon usage and utilization of C and G relative to gene expression (as measured by EST redundance) were studied, and preferences for utilization of C and CpG doublets were noted among the P. tricornutum EST coding sequences.

Animals↗

PRINTS--a protein motif fingerprint database.

The PRINTS database of protein 'fingerprints' is described. Fingerprints comprise sets of motifs excised from conserved regions of sequence alignments, their diagnostic power or potency being refined by iterative database scanning (in this case the OWL composite sequence database). Generally, the motifs do not overlap, but are separated along a sequence, though they may be contiguous in 3-D space. The use of groups of independent, linearly or spatially separate motifs allows particular protein folds and functionalities to be characterized more flexibly and powerfully than conventional single-component patterns or regular expressions. The current version of the database (4.0) contains 150 entries (encoding > 700 motifs), covering a wide range of globular and membrane proteins, modular polypeptides and so on. The growth of the database is influenced by a number of factors, e.g. the use of multiple motifs, the maximization of sequence information through iterative database scanning and the fact that the database searched is a large composite. The information contained within PRINTS is distinct from but complementary to the single consensus expressions stored in the widely used PROSITE dictionary of patterns.

Amino Acid Sequence↗

Efficiency of database search for identification of mutated and modified proteins via mass spectrometry.

Although protein identification by matching tandem mass spectra (MS/MS) against protein databases is a widespread tool in mass spectrometry, the question about reliability of such searches remains open. Absence of rigorous significance scores in MS/MS database search makes it difficult to discard random database hits and may lead to erroneous protein identification, particularly in the case of mutated or post-translationally modified peptides. This problem is especially important for high-throughput MS/MS projects when the possibility of expert analysis is limited. Thus, algorithms that sort out reliable database hits from unreliable ones and identify mutated and modified peptides are sought. Most MS/MS database search algorithms rely on variations of the Shared Peaks Count approach that scores pairs of spectra by the peaks (masses) they have in common. Although this approach proved to be useful, it has a high error rate in identification of mutated and modified peptides. We describe new MS/MS database search tools, MS-CONVOLUTION and MS-ALIGNMENT, which implement the spectral convolution and spectral alignment approaches to peptide identification. We further analyze these approaches to identification of modified peptides and demonstrate their advantages over the Shared Peaks Count. We also use the spectral alignment approach as a filter in a new database search algorithm that reliably identifies peptides differing by up to two mutations/modifications from a peptide in a database.

Algorithms↗