PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Annotated proteome of a human T-cell lymphoma.

As the reliable identification of proteins by tandem mass spectrometry becomes increasingly common, the full characterization of large data sets of proteins remains a difficult challenge. Our goal was to survey the proteome of a human T-cell lymphoma-derived cell line in a single set of experiments and present an automated method for the annotation of lists of proteins. A downstream application of these data includes the identification of novel pathogenetic and candidate diagnostic markers of T-cell lymphoma. Total protein isolated from cytoplasmic, membrane, and nuclear fractions of the SUDHL-1 T-cell lymphoma cell line was resolved by SDS-PAGE, and the entire gel lanes digested and analyzed by tandem mass spectrometry. Acquired data files were searched against the UniProt protein database using the SEQUEST algorithm. Search results for each subcellular fraction were analyzed using INTERACT and ProteinProphet. All protein identifications with an error rate of less than 10% were directly exported into excel and analyzed using GOMiner (NIH/NCI). The Gene ontology molecular function and cell location data were summarized for the identified proteins and results exported as user-interactive directed acyclic graphs. A total of 1105 unique proteins were identified and fully annotated, including numerous proteins that had not been previously characterized in lymphoma, in functional categories such as cell adhesion, migration, signaling, and stress response. This study demonstrates the utility of currently available bioinformatics tools for the robust identification and annotation of large numbers of proteins in a batchwise fashion.

Algorithms↗

Identification of human whole saliva protein components using proteomics.

The determination of salivary biomarkers as a means of monitoring general health and for the early diagnosis of disease is of increasing interest in clinical research. Based on the linkage between salivary proteins and systemic diseases, the aim of this work was the identification of saliva proteins using proteomics. Salivary proteins were separated using two-dimensional (2-D) gel electrophoresis over a pH range between 3-10, digested, and then analyzed by matrix assisted laser desorption/ionization-time of flight (MALDI-TOF)-TOF mass spectrometry (MS) and tandem mass spectrometry (MS/MS). Proteins were identified using automated MS and MS/MS data acquisition. The resulting data were searched against a protein database using an internal Mascot search routine. Ninety spots give identifications with high statistical reliability. Of the identified proteins, 11 were separated and identified in saliva for the first time using proteomics tools. Moreover, three proteins that have not been previously identified in saliva, PLUNC, cystatin A, and cystatin B were identified.

Cystatin B↗

The Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics that supports genomic and proteomic research and scientific discovery. PIR maintains the Protein Sequence Database (PSD), an annotated protein database containing over 283 000 sequences covering the entire taxonomic range. Family classification is used for sensitive identification, consistent annotation, and detection of annotation errors. The superfamily curation defines signature domain architecture and categorizes memberships to improve automated classification. To increase the amount of experimental annotation, the PIR has developed a bibliography system for literature searching, mapping, and user submission, and has conducted retrospective attribution of citations for experimental features. PIR also maintains NREF, a non-redundant reference database, and iProClass, an integrated database of protein family, function, and structure information. PIR-NREF provides a timely and comprehensive collection of protein sequences, currently consisting of more than 1 000 000 entries from PIR-PSD, SWISS-PROT, TrEMBL, RefSeq, GenPept, and PDB. The PIR web site (http://pir.georgetown.edu) connects data analysis tools to underlying databases for information retrieval and knowledge discovery, with functionalities for interactive queries, combinations of sequence and text searches, and sorting and visual exploration of search results. The FTP site provides free download for PSD and NREF biweekly releases and auxiliary databases and files.

Amino Acid Sequence↗

A reference map of a human pituitary adenoma proteome.

In order to compare the proteomes from different cell types of pituitary adenomas for our long-term goal to clarify the molecular mechanisms that participate in the formation of pituitary adenoma, and to detect any tumor-related marker for an "early-stage" diagnosis, the two-dimensional gel electrophoresis (2-DE) reference map of a pituitary adenoma tissue proteome is described here. A vertical, two-dimensional (2-D) polyacrylamide gel electrophoresis system and PDQuest image analysis software have been used to provide a high level of between-gel reproducibility and to accurately array each protein expressed in a pituitary adenoma tissue. Mass spectrometry (matrix-assisted laser desorption/ionization-time of flight MALDI-TOF and liquid chromatography-electrospray ionization-quadrupole-ion trap LC-ESI-Q-IT) and protein databases were used to characterize each protein in the 2-D gel. The results demonstrate that a good reproducibility of the 2-D gel pattern was attained. The position deviation of matched spots among four 2-D gels was 1.95 +/- 0.45 mm in the isoelectric focusing direction, and 1.70 +/- 0.53 mm in the sodium dodecyl sulfate-polyacrylamide gel electrophoresis direction. A total of ca. 1000 protein spots were separated by 2-DE, and 135 protein spots that represent 111 proteins were characterized with mass spectrometry (96 spots for MALDI-TOF, 39 spots for LC-ESI-Q-IT). The characterized proteins include pituitary hormones, cellular signals, enzymes, cellular-defense proteins, cell-structure proteins, transport proteins, etc. Those proteins were located in the cytoplasmic, cellular membrane, mitochondrial, endoplasmic reticulum, nuclear, ribonucleosome, extracellular fractions, or were secreted in plasma, etc. Those identified proteins contribute to a functional profile of the pituitary adenoma proteome. These data will be used to expand the proteome database of the human pituitary, which can be accessed in the website http://www.utmem.edu /proteomics.

Adenoma↗

Including mutations from conceptually translated expressed sequence tags into orthologous proteins improves the preliminary assignment of peptide mass fingerprints on non-model genomes.

In order to improve protein assignment from peptide mass fingerprints (PMF) in species with incompletely sequenced genomes, the genus-specific mutations deduced from Expressed sequence tag (EST) sequences were included in the complete reading frames of orthologous proteins, resulting in a new searchable in silico protein database. Using this method in tests on four plant species, the MOWSE score of at least 20% more proteins was improved compared to conventional approaches on crude, total proteins, for middle-sized EST projects. Larger contigs are assembled in more important EST projects and this improves the conventional assignment of the most abundant proteins. However, contigs from minor transcripts remain shorter and the assignment of less abundant proteins, such as those isolated following subcellular fractionment, is improved by searching orthologue-EST conceptual chimeras with the PMF spectra. This strategy may be utilized as a tool to identify potential PMF matches that can be then verified by other experimental approaches (tandem mass spectrometry) to ensure the EST matched chimera identification is accurate.

Amino Acid Sequence↗

Protein ranking by semi-supervised network propagation.

BACKGROUND: Biologists regularly search DNA or protein databases for sequences that share an evolutionary or functional relationship with a given query sequence. Traditional search methods, such as BLAST and PSI-BLAST, focus on detecting statistically significant pairwise sequence alignments and often miss more subtle sequence similarity. Recent work in the machine learning community has shown that exploiting the global structure of the network defined by these pairwise similarities can help detect more remote relationships than a purely local measure. METHODS: We review RankProp, a ranking algorithm that exploits the global network structure of similarity relationships among proteins in a database by performing a diffusion operation on a protein similarity network with weighted edges. The original RankProp algorithm is unsupervised. Here, we describe a semi-supervised version of the algorithm that uses labeled examples. Three possible ways of incorporating label information are considered: (i) as a validation set for model selection, (ii) to learn a new network, by choosing which transfer function to use for a given query, and (iii) to estimate edge weights, which measure the probability of inferring structural similarity. RESULTS: Benchmarked on a human-curated database of protein structures, the original RankProp algorithm provides significant improvement over local network search algorithms such as PSI-BLAST. Furthermore, we show here that labeled data can be used to learn a network without any need for estimating parameters of the transfer function, and that diffusion on this learned network produces better results than the original RankProp algorithm with a fixed network. CONCLUSION: In order to gain maximal information from a network, labeled and unlabeled data should be used to extract both local and global structure.

Algorithms↗

Sequence and hydropathy profile analysis of two classes of secondary transporters.

A structural class in the MemGen classification of membrane proteins is a set of evolutionary related proteins sharing a similar global fold. A structural class contains both closely related pairs of proteins for which homology is clear from sequence comparison and very distantly related pairs, for which it is not possible to establish homology based on sequence similarity alone. In the latter case the evolutionary link is based on hydropathy profile analysis. Here, we use these evolutionary related sets of proteins to analyze the relationship between E-values in BLAST searches, sequence similarities in multiple sequence alignments and structural similarities in hydropathy profile analyses. Two structural classes of secondary transporters termed ST[3], which includes the Ion Transporter (IT) superfamily and ST[4], which includes the DAACS family (TC# 2.A.23) were extracted from the NCBI protein database. ST[3] contains 2051 unique sequences distributed over 32 families and 59 subfamilies. ST[4] is a smaller class containing 399 unique sequences distributed over 2 families and 7 subfamilies. One subfamily in ST[4] contains a new class of binding protein dependent secondary transporters. Comparison of the averaged hydropathy profiles of the subfamilies in ST[3] and ST[4] revealed that the two classes represent different folds. Divergence of the sequences in ST[4] is much smaller than observed in ST[3], suggesting different constraints on the proteins during evolution. Analysis of the correlation between the evolutionary relationship of pairs of proteins in a class and the BLAST E-value revealed that: (i) the BLAST algorithm is unable to pick up the majority of the links between proteins in structural class ST[3], (ii) "low complexity filtering" and "composition based statistics" improve the specificity, but strongly reduce the sensitivity of BLAST searches for distantly related proteins, indicating that these filters are too stringent for the proteins analyzed, and (iii) the E-value cut-off, which may be used to evaluate evolutionary significance of a hit in a BLAST search is very different for the two structural classes of membrane proteins.

Algorithms↗

SPIDER: software for protein identification from sequence tags with de novo sequencing error.

For the identification of novel proteins using MS/MS, de novo sequencing software computes one or several possible amino acid sequences (called sequence tags) for each MS/MS spectrum. Those tags are then used to match, accounting amino acid mutations, the sequences in a protein database. If the de novo sequencing gives correct tags, the homologs of the proteins can be identified by this approach and software such as MS-BLAST is available for the matching. However, de novo sequencing very often gives only partially correct tags. The most common error is that a segment of amino acids is replaced by another segment with approximately the same masses. We developed a new efficient algorithm to match sequence tags with errors to database sequences for the purpose of protein and peptide identification. A software package, SPIDER, was developed and made available on Internet for free public use. This paper describes the algorithms and features of the SPIDER software.

Algorithms↗

Do structurally similar ligands bind in a similar fashion?

The scope of the current work is to investigate whether structurally similar ligands bind in a similar fashion by exhaustively analyzing experimental data from the protein database (PDB). The complete PDB was searched for pairs of structurally similar ligands binding to the same biological target. The binding sites of the pairs of proteins complexing structurally similar ligands were found to differ in 83% of the cases. The most recurrent structural change among the pairs involves different water molecule architecture. Side-chain movements are observed in half of the pairs, whereas backbone movements rarely occurred. However, two structurally similar ligands generally confirm a high degree of structural conservation. That is, a majority of the ligand pairs occupy the same region in the binding sites, providing support for the use of shape matching in the drug design process. We allow ourselves to draw general conclusions because our data set consists of ligands with drug-like physicochemical properties complexed to a broad spectrum of different protein classes.

Crystallography, X-Ray↗

Protein three-dimensional structural databases: domains, structurally aligned homologues and superfamilies.

This paper reports the availability of a database of protein structural domains (DDBASE), an alignment database of homologous proteins (HOMSTRAD) and a database of structurally aligned superfamilies (CAMPASS) on the World Wide Web (WWW). DDBASE contains information on the organization of structural domains and their boundaries; it includes only one representative domain from each of the homologous families. This database has been derived by identifying the presence of structural domains in proteins on the basis of inter-secondary structural distances using the program DIAL [Sowdhamini & Blundell (1995), Protein Sci. 4, 506-520]. The alignment of proteins in superfamilies has been performed on the basis of the structural features and relationships of individual residues using the program COMPARER [Sali & Blundell (1990), J. Mol. Biol. 212, 403-428]. The alignment databases contain information on the conserved structural features in homologous proteins and those belonging to superfamilies. Available data include the sequence alignments in structure-annotated formats and the provision for viewing superposed structures of proteins using a graphical interface. Such information, which is freely accessible on the WWW, should be of value to crystallographers in the comparison of newly determined protein structures with previously identified protein domains or existing families.

Amino Acid Sequence↗

Defining parameters for homology-tolerant database searching.

De novo interpretation of tandem mass spectrometry (MS/MS) spectra provides sequences for searching protein databases when limited sequence information is present in the database. Our objective was to define a strategy for this type of homology-tolerant database search. Homology searches, using MS-Homology software, were conducted with 20, 10, or 5 of the most abundant peptides from 9 proteins, based either on precursor trigger intensity or on total ion current, and allowing for 50%, 30%, or 10% mismatch in the search. Protein scores were corrected by subtracting a threshold score that was calculated from random peptides. The highest (p < .01) corrected protein scores (i.e., above the threshold) were obtained by submitting 20 peptides and allowing 30% mismatch. Using these criteria, protein identification based on ion mass searching using MS/MS data (i.e., Mascot) was compared with that obtained using homology search. The highest-ranking protein was the same using Mascot, homology search using the 20 most intense peptides, or homology search using all peptides, for 63.4% of 112 spots from two-dimensional polyacrylamide gel electrophoresis gels. For these proteins, the percent coverage was greatest using Mascot compared with the use of all or just the 20 most intense peptides in a homology search (25.1%, 18.3%, and 10.6%, respectively). Finally, 35% of de novo sequences completely matched the corresponding known amino acid sequence of the matching peptide. This percentage increased when the search was limited to the 20 most intense peptides (44.0%). After identifying the protein using MS-Homology, a peptide mass search may increase the percent coverage of the protein identified.

Amino Acid Sequence↗

BEAUTY-X: enhanced BLAST searches for DNA queries.

UNLABELLED: BEAUTY (BLAST Enhanced Alignment Utility) is an enhanced version of the BLAST database search tool that facilitates identification of the functions of matched sequences. Three recent improvements to the BEAUTY program described here make the enhanced output (1) available for DNA queries, (2) available for searches of any protein database, and (3) more up-to-date, with periodic updates of the domain information. AVAILABILITY: BEAUTY searches of the NCBI and EMBL non-redundant protein sequence databases are available from the BCM Search Launcher Web pages (http://gc.bcm.tmc. edu:8088/search-launcher/launcher.html). BEAUTY Post-Processing of submitted search results is available using the BCM Search Launcher Batch Client (version 2.6) (ftp://gc.bcm.tmc. edu/pub/software/search-launcher/). SUPPLEMENTARY INFORMATION: Example figures are available at http://dot.bcm.tmc. edu:9331/papers/beautypp.html CONTACT: (kworley,culpep)@bcm.tmc.edu

Amino Acid Sequence↗

Rapid enrichment of bioactive milk proteins and iterative, consolidated protein identification by multidimensional protein identification technology.

Direct injection of complex protein mixtures, e.g. those derived from crude biological fluids, is often incompatible with conventional LC supports, because of column clogging and rapid deterioration of chromatographic performance. In this paper, we report the use of restricted access media to rapidly enrich and fractionate human breast milk. This resin, combining size exclusion and anion exchange functionalities, yielded a fraction enriched in soluble CD14 and showing specific sCD14-dependant activity. This fraction was split into five aliquots, which were individually characterized using multidimensional protein identification technology. Reproducibility of the results was addressed by analysing and comparing five datasets using different protein identification tools available within the Sequest software. Furthermore, a comparison of three major releases of the Ensembl human protein database was performed to examine the effect of database updates on our results. We report here the benefit of repeated analysis of aliquots of the same fraction: first to increase the confidence in peptide identification by repeated confirmation in several aliquots; and second to assess experimental reproducibility. We demonstrate furthermore the effect of database modifications on the results and the importance of constantly re-analysing data with new releases to keep them consistent and up to date with the latest protein identities and predictions available.

Amino Acid Sequence↗

Coverage of whole proteome by structural genomics observed through protein homology modeling database.

We have been developing FAMSBASE, a protein homology-modeling database of whole ORFs predicted from genome sequences. The latest update of FAMSBASE ( http://daisy.nagahama-i-bio.ac.jp/Famsbase/ ), which is based on the protein three-dimensional (3D) structures released by November 2003, contains modeled 3D structures for 368,724 open reading frames (ORFs) derived from genomes of 276 species, namely 17 archaebacterial, 130 eubacterial, 18 eukaryotic and 111 phage genomes. Those 276 genomes are predicted to have 734,193 ORFs in total and the current FAMSBASE contains protein 3D structure of approximately 50% of the ORF products. However, cases that a modeled 3D structure covers the whole part of an ORF product are rare. When portion of an ORF with 3D structure is compared in three kingdoms of life, in archaebacteria and eubacteria, approximately 60% of the ORFs have modeled 3D structures covering almost the entire amino acid sequences, however, the percentage falls to about 30% in eukaryotes. When annual differences in the number of ORFs with modeled 3D structure are calculated, the fraction of modeled 3D structures of soluble protein for archaebacteria is increased by 5%, and that for eubacteria by 7% in the last 3 years. Assuming that this rate would be maintained and that determination of 3D structures for predicted disordered regions is unattainable, whole soluble protein model structures of prokaryotes without the putative disordered regions will be in hand within 15 years. For eukaryotic proteins, they will be in hand within 25 years. The 3D structures we will have at those times are not the 3D structure of the entire proteins encoded in single ORFs, but the 3D structures of separate structural domains. Measuring or predicting spatial arrangements of structural domains in an ORF will then be a coming issue of structural genomics.

Amino Acid Sequence↗

IdentiCS--identification of coding sequence and in silico reconstruction of the metabolic network directly from unannotated low-coverage bacterial genome sequence.

BACKGROUND: A necessary step for a genome level analysis of the cellular metabolism is the in silico reconstruction of the metabolic network from genome sequences. The available methods are mainly based on the annotation of genome sequences including two successive steps, the prediction of coding sequences (CDS) and their function assignment. The annotation process takes time. The available methods often encounter difficulties when dealing with unfinished error-containing genomic sequence. RESULTS: In this work a fast method is proposed to use unannotated genome sequence for predicting CDSs and for an in silico reconstruction of metabolic networks. Instead of using predicted genes or CDSs to query public databases, entries from public DNA or protein databases are used as queries to search a local database of the unannotated genome sequence to predict CDSs. Functions are assigned to the predicted CDSs simultaneously. The well-annotated genome of Salmonella typhimurium LT2 is used as an example to demonstrate the applicability of the method. 97.7% of the CDSs in the original annotation are correctly identified. The use of SWISS-PROT-TrEMBL databases resulted in an identification of 98.9% of CDSs that have EC-numbers in the published annotation. Furthermore, two versions of sequences of the bacterium Klebsiella pneumoniae with different genome coverage (3.9 and 7.9 fold, respectively) are examined. The results suggest that a 3.9-fold coverage of the bacterial genome could be sufficiently used for the in silico reconstruction of the metabolic network. Compared to other gene finding methods such as CRITICA our method is more suitable for exploiting sequences of low genome coverage. Based on the new method, a program called IdentiCS (Identification of Coding Sequences from Unfinished Genome Sequences) is delivered that combines the identification of CDSs with the reconstruction, comparison and visualization of metabolic networks (free to download at http://genome.gbf.de/bioinformatics/index.html). CONCLUSIONS: The reversed querying process and the program IdentiCS allow a fast and adequate prediction protein coding sequences and reconstruction of the potential metabolic network from low coverage genome sequences of bacteria. The new method can accelerate the use of genomic data for studying cellular metabolism.

Base Sequence↗

Spermatocytes and round spermatids of rat testis: protein patterns.

Spermatogenesis is a process in the testis that involves meiotic cell division and spermiogenesis. The mechanisms of regulation and its associated proteins are mostly unknown. This publication shows the two-dimensional (2-D) gel electrophoresis protein map obtained from rat testis using nonlinear 3.5-10 immobilized pH gradients for the first-dimensional separation. Eighteen proteins were successfully identified in the SWISS-PROT protein database using amino acid analysis of proteins recovered from polyvinylidene difluoride (PVDF) membranes and verified for one of them by comparison with Anderson's rat liver reference map. Fourteen new polypeptides were identified and four were previously known. Two of these new proteins were closely related to the spermatogenetic process. T-complex protein 1 is expressed in large amounts in germ cells. Androgen-dependent sperm-coating glycoprotein is secreted by epididymal cells. In order to detect changes in protein expression during meiosis and spermiogenesis, spermatocytes and round spermatid cell populations were purified by centrifugal elutriation and compared. In this way several proteins not found in the spermatocyte 2-D images could be high-lighted. The sperm-coating glycoprotein was thus shown to be present in large amounts in round spermatids.

Adolescent↗

Identification of Francisella tularensis genes encoding exported membrane-associated proteins using TnphoA mutagenesis of a genomic library.

Francisella tularensis, the causative agent of tularemia, is a highly infectious pathogen of humans and animals, yet little is known about the surface proteins of this organism that mediate mechanisms of pathogenicity. lambdaTnphoA was used to generate random alkaline phosphatase gene fusions in a F. tularensis subsp. tularensis (strain Schu S4) genomic library to identify genes encoding exported extracytoplasmic proteins. Eleven genes encoding membrane-associated proteins were identified by this method and their respective signal peptides were characterized. Three of the genes encoded conserved 'housekeeping' enzymes, while the other eight genes were unique to F. tularensis, encoding proteins with molecular masses ranging from 11 to 78kDa as deduced from the amino acid sequences. Two genes putatively encoded lipoproteins based on the presence of characteristic signal peptidase II cleavage sites. Four selected proteins were found associated with outer membranes from Schu S4 and LVS strains by Western blotting. Indirect immunofluorescence of strain Schu S4 cells also showed evidence of protein localization to the outer membrane. Protein database searches produced significant alignments with proteins from other bacteria involved in carbohydrate transport, lipid metabolism, and cell envelope biogenesis, thereby providing clues for putative functions. These findings demonstrated that TnphoA mutagenesis can be used in conjunction with F. tularensis genome sequence data to provide a foundation for studies to identify and define cellular surface protein virulence factors of this pathogen.

Alkaline Phosphatase↗

Yeast genomic expression studies using DNA microarrays.

The exploration and characterization of yeast genomic expression programs is providing a wealth of information about yeast biology, as well as other organisms. The intriguing biology of yeast species invites characterization of genomic expression patterns to illuminate the details of cellular physiology. In addition to its value as an interesting organism, yeast maintains its role as an excellent model in which to characterize genomic expression programs. Microarray studies are quickly spreading to plant, animal, and microbial organisms that remain in the early stages of characterization. The extensive knowledge of yeast biology, as well as the relative ease with which yeast studies can be performed and controlled, facilitates interpretation of the genomic expression data. Importantly, existing information about yeast biology, including functional annotations for each gene, is captured and efficiently presented in databases such as the Saccharomyces Genome Database (SGD), the Munich Information Center Yeast Genome Database (MIPS), the Yeast and Pombe Protein Databases (YPD and PPD, respectively), and others. A number of databases also allow the exploration of published genomic expression studies, including the "Expression Connection" at SGD and the Microarray Global Viewer (yMGV) organized by Marc et al. Consulting these databases to retrieve known details about gene function and regulation vastly facilitates interpretation of the genomic expression data, allowing biological hypotheses to be formulated and tested. These hypotheses can be applied to other organisms that may execute genomic expression programs similar to those seen in yeast. Furthermore, as more genomic expression studies in multiple organisms emerge, large-scale data comparisons can be conducted, within and across organisms. Incorporating the results of yeast studies into such comparisons is certain to increase our understanding about the function, regulation, and evolution of genomic expression programs.

Carbocyanines↗