PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Adult midgut expressed sequence tags from the tsetse fly Glossina morsitans morsitans and expression analysis of putative immune response genes.

BACKGROUND: Tsetse flies transmit African trypanosomiasis leading to half a million cases annually. Trypanosomiasis in animals (nagana) remains a massive brake on African agricultural development. While trypanosome biology is widely studied, knowledge of tsetse flies is very limited, particularly at the molecular level. This is a serious impediment to investigations of tsetse-trypanosome interactions. We have undertaken an expressed sequence tag (EST) project on the adult tsetse midgut, the major organ system for establishment and early development of trypanosomes. RESULTS: A total of 21,427 ESTs were produced from the midgut of adult Glossina morsitans morsitans and grouped into 8,876 clusters or singletons potentially representing unique genes. Putative functions were ascribed to 4,035 of these by homology. Of these, a remarkable 3,884 had their most significant matches in the Drosophila protein database. We selected 68 genes with putative immune-related functions, macroarrayed them and determined their expression profiles following bacterial or trypanosome challenge. In both infections many genes are downregulated, suggesting a malaise response in the midgut. Trypanosome and bacterial challenge result in upregulation of different genes, suggesting that different recognition pathways are involved in the two responses. The most notable block of genes upregulated in response to trypanosome challenge are a series of Toll and Imd genes and a series of genes involved in oxidative stress responses. CONCLUSIONS: The project increases the number of known Glossina genes by two orders of magnitude. Identification of putative immunity genes and their preliminary characterization provides a resource for the experimental dissection of tsetse-trypanosome interactions.

Aging↗

Genome sequencing and annotation: an overview.

Many microbial genome sequences have been determined, and more new genome projects are ongoing. Shotgun sequencing of randomly cloned short pieces of genomic DNA can provide a simple way of determining whole genome sequences. This process requires sequencing of many fragments, compilation of the separate sequences into one contiguous sequence, and careful editing of the assembled sequence. The genes present on the microbial genome are then predicted using clues derived from typical gene features, such as codon usage, ribosomal binding sequences, and bacterial initiation codons. Function of genes is predicted by homology searches performed against either public or well-established protein databases. This chapter discusses each of these stages in a genome-sequencing project.

Amino Acid Sequence↗

Changes of chondrocyte metabolism in vitro: an approach by proteomic analysis.

Changes in chondrocyte metabolism in vitro using different support systems and under different culture conditions were studied with a proteomic approach. Qualitative and quantitative modifications in the synthesis of chondrocyte proteins were investigated using two-dimensional (2D) gel electrophoresis. This technique provided a simple way to visualize the most abundant chondrocyte proteins. Proteins were identified after in-gel proteolysis with trypsin and matrix-assisted laser desorption ionization-time of flight mass spectrometry, using peptide mass fingerprinting. Tryptic peptide masses were measured and matched against a computer-generated list from the simulated trypsin proteolysis of a protein database (SwissProt).

Animals↗

Peptide mass fingerprinting: identification of proteins by MALDI-TOF.

MALDI-TOF peptide mass fingerprinting (PMF) is the fastest and cheapest method of protein identification; the studied genome is sequenced and annotated, and the protein is amenable to separation and detection in 2D gel electrophoresis. In plant proteomics there are two main difficulties: few plant genomes are sequenced, and major contaminants are non-plant specific. This chapter describes the classical "bottom-up" method (i.e., from peptide to protein identification) of gel cutting, in-gel digestion, peptide recovery and purification, MALDI-TOF mass spectrometry, and critical survey of protein database queries.

Acrylic Resins↗

Interpretation of collision-induced fragmentation tandem mass spectra of posttranslationally modified peptides.

Tandem collision-induced dissociation (CID) mass spectrometry (MS) provides a sensitive means of analyzing the amino acid sequence of peptides. Modern MS instrumentation is capable of rapidly generating many thousands of tandem mass spectra, and protein database search engines have been developed to cope with this avalanche of data. In most studies, there is a schism between discarding perfectly valid data and including nonsensical peptide identifications--this is currently a major bottleneck in data analysis and it calls for manual evaluation of the data. Especially for posttranslationally modified peptides, there is a need for manual validation of the data because search algorithms seldom have been optimized for the identification of modified peptides and because there are many pitfalls for the unwary. This chapter describes some of the issues that should be considered when interpreting and validating low-energy CID tandem mass spectra and gives some useful tables to aid this process.

Amino Acid Sequence↗

Evaluation of the information content in infrared spectra for protein secondary structure determination.

Fourier-transform infrared spectroscopy is a method of choice for the experimental determination of protein secondary structure. Numerous approaches have been developed during the past 15 years. A critical parameter that has not been taken into account systematically is the selection of the wavenumbers used for building the mathematical models used for structure prediction. The high quality of the current Fourier-transform infrared spectrometers makes the absorbance at every single wavenumber a valid and almost noiseless type of information. We address here the question of the amount of independent information present in the infrared spectra of proteins for the prediction of the different secondary structure contents. It appears that, at most, the absorbance at three distinct frequencies of the spectra contain all the nonredundant information that can be related to one secondary structure content. The ascending stepwise method proposed here identifies the relevance of each wavenumber of the infrared spectrum for the prediction of a given secondary structure and yields a particularly simple method for computing the secondary structure content. Using the 50-protein database built beforehand to contain as little fold redundancy as possible, the standard error of prediction in cross-validation is 5.5% for the alpha-helix, 6.6% for the beta-sheet, and 3.4% for the beta-turn.

Algorithms↗

BioParser: a tool for processing of sequence similarity analysis reports.

UNLABELLED: The widely used programs BLAST (in this article, 'BLAST' includes both the National Center for Biotechnology Information [NCBI] BLAST and the Washington University version WU BLAST) and FASTA for similarity searches in nucleotide and protein databases usually result in copious output. However, when large query sets are used, human inspection rapidly becomes impractical. BioParser is a Perl program for parsing BLAST and FASTA reports. Making extensive use of the BioPerl toolkit, the program filters, stores and returns components of these reports in either ASCII or HTML format. BioParser is also capable of automatically feeding a local MySQL database with the parsed information, allowing subsequent filtering of hits and/or alignments with specific attributes. For this reason, BioParser is a valuable tool for large-scale similarity analyses by improving the access to the information present in BLAST or FASTA reports, facilitating extraction of useful information of large sets of sequence alignments, and allowing for easy handling and processing of the data. AVAILABILITY: BioParser is licensed under the Creative Commons Attribution-NonCommercial-NoDerivs 2.0 license terms (http://creativecommons.org/licenses/by-nc-nd/2.0/) and is available upon request. Additional information can be found at the BioParser website (http://www.dbbm.fiocruz.br/BioParser.html).

Amino Acid Sequence↗

Discrimination of non-protein-coding transcripts from protein-coding mRNA.

Several recent studies indicate that mammals and other organisms produce large numbers of RNA transcripts that do not correspond to known genes. It has been suggested that these transcripts do not encode proteins, but may instead function as RNAs. However, discrimination of coding and non-coding transcripts is not straightforward, and different laboratories have used different methods, whose ability to perform this discrimination is unclear. In this study, we examine ten bioinformatic methods that assess protein-coding potential and compare their ability and congruency in the discrimination of non-coding from coding sequences, based on four underlying principles: open reading frame size, sequence similarity to known proteins or protein domains, statistical models of protein-coding sequence, and synonymous versus non-synonymous substitution rates. Despite these different approaches, the methods show broad concordance, suggesting that coding and non-coding transcripts can, in general, be reliably discriminated, and that many of the recently discovered extra-genic transcripts are indeed non-coding. Comparison of the methods indicates reasons for unreliable predictions, and approaches to increase confidence further. Conversely and surprisingly, our analyses also provide evidence that as much as approximately 10% of entries in the manually curated protein database Swiss-Prot are erroneous translations of actually non-coding transcripts.

Algorithms↗

Resistance gene analogues of Arabidopsis thaliana: recognition by structure.

Following completion of Arabidopsis thaliana sequencing projects, multiple resistance gene analogues (RGAs) have been identified. In this work a review of the current state of knowledge available in protein databases and scientific articles is presented. Putative resistance genes were identified by using BLAST searches as well as HMM fingerprints (the latter to infer existence of characteristic domains). The representation of all five classes of putative resistance genes in Col-0 ecotype was examined, along with the statistics on RGAs present on all five chromosomes of Arabidopsis thaliana.

Arabidopsis↗

Incremental generation of summarized clustering hierarchy for protein family analysis.

MOTIVATION: Protein sequence clustering has been widely exploited to facilitate in-depth analysis of protein functions and families. For some applications of protein sequence clustering, it is highly desirable that a hierarchical structure, also referred to as dendrogram, which shows how proteins are clustered at various levels, is generated. However, as the sizes of contemporary protein databases continue to grow at rapid rates, it is of great interest to develop some summarization mechanisms so that the users can browse the dendrogram and/or search for the desired information more effectively. RESULTS: In this paper, the design of a novel incremental clustering algorithm aimed at generating summarized dendrograms for analysis of protein databases is described. The proposed incremental clustering algorithm employs a statistics-based model to summarize the distributions of the similarity scores among the proteins in the database and to control formation of clusters. Experimental results reveal that, due to the summarization mechanism incorporated, the proposed incremental clustering algorithm offers the users highly concise dendrograms for analysis of protein clusters with biological significance. Another distinction of the proposed algorithm is its incremental nature. As the sizes of the contemporary protein databases continue to grow at fast rates, due to the concern of efficiency, it is desirable that cluster analysis of a protein database can be carried out incrementally, when the protein database is updated. Experimental results with the Swiss-Prot protein database reveal that the time complexity for carrying out incremental clustering with k new proteins added into the database containing n proteins is O(n2betalogn), where beta congruent with 0.865, provided that k << n. AVAILABILITY: The Linux executable is available on the following supplementary page.

Algorithms↗

3-D lookup: fast protein structure database searches at 90% reliability.

There are far fewer classes of three-dimensional protein folds than sequence families but the problem of detecting three-dimensional similarities is NP-complete. We present a novel heuristic for identifying 3-D similarities between a query structure and the database of known protein structures. Many methods for structure alignment use a bottom-up approach, identifying first local matches and then solving a combinatorial problem in building up larger clusters of matching substructures. Here, the top-down approach is to start with the global comparison and select a rough superimposition using a fast 3-D lookup of secondary structure motifs. The superimposition is then extended to an alignment of C alpha atoms by an iterative dynamic programming step. An all-against-all comparison of 385 representative proteins (150,000 pair comparisons) took 1 day of computer time on a single R8000 processor. In other words, one query structure is scanned against the database in a matter of minutes. The method is rated at 90% reliability at capturing statistically significant similarities. It is useful as a rapid preprocessor to a comprehensive protein structure database search system.

Amino Acid Sequence↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural three dimensional (3-D) and sequence one dimensional(1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in Swissprot using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 27% of all Swissprot-stored sequences.

Amino Acid Sequence↗

CellCircuits: a database of protein network models.

CellCircuits (http://www.cellcircuits.org) is an open-access database of molecular network models, designed to bridge the gap between databases of individual pairwise molecular interactions and databases of validated pathways. CellCircuits captures the output from an increasing number of approaches that screen molecular interaction networks to identify functional subnetworks, based on their correspondence with expression or phenotypic data, their internal structure or their conservation across species. This initial release catalogs 2019 computationally derived models drawn from 11 journal articles and spanning five organisms (yeast, worm, fly, Plasmodium falciparum and human). Models are available either as images or in machine-readable formats and can be queried by the names of proteins they contain or by their enriched biological functions. We envision CellCircuits as a clearinghouse in which theorists may distribute or revise models in need of validation and experimentalists may search for models or specific hypotheses relevant to their interests. We demonstrate how such a repository of network models is a novel systems biology resource by performing several meta-analyses not currently possible with existing databases.

Animals↗

ProTherm, version 2.0: thermodynamic database for proteins and mutants.

ProTherm 2.0 is the second release of the Thermo-dynamic Database for Proteins and Mutants that includes numerical data for several thermodynamic parameters, structural information, experimental methods and conditions, functional and literature information. The present release contains >5500 entries, an approximately 67% increase over the previous version. In addition, we have included information about reversibility of data, details about buffer and ion concentrations and the surrounding residues in space for all mutants. A WWW interface enables users to search data based on various conditions with different sorting options for outputs. Further, ProTherm has links with other structural and literature databases, and the mutation sites and surrounding residues are automatically mapped on the structures and can be directly viewed through 3DinSight developed in our laboratory. The ProTherm database is freely available through the WWW at http://www.rtc.riken.go.jp/protherm.html

Databases, Factual↗

Frequency of gaps observed in a structurally aligned protein pair database suggests a simple gap penalty function.

Gap penalty is an important component of the scoring scheme that is needed when searching for homologous proteins and for accurate alignment of protein sequences. Most homology search and sequence alignment algorithms employ a heuristic 'affine gap penalty' scheme q + r x n, in which q is the penalty for opening a gap, r the penalty for extending it and n the gap length. In order to devise a more rational scoring scheme, we examined the pattern of gaps that occur in a database of structurally aligned protein domain pairs. We find that the logarithm of the frequency of gaps varies linearly with the length of the gap, but with a break at a gap of length 3, and is well approximated by two linear regression lines with R2 values of 1.0 and 0.99. The bilinear behavior is retained when gaps are categorized by secondary structures of the two residues flanking the gap. Similar results were obtained when another, totally independent, structurally aligned protein pair database was used. These results suggest a modification of the affine gap penalty function.

Computational Biology↗

Flexible Structural Neighborhood--a database of protein structural similarities and alignments.

Protein structures are flexible, changing their shapes not only upon substrate binding, but also during evolution as a collective effect of mutations, deletions and insertions. A new generation of protein structure comparison algorithms allows for such flexibility; they go beyond identifying the largest common part between two proteins and find hinge regions and patterns of flexibility in protein families. Here we present a Flexible Structural Neighborhood (FSN), a database of structural neighbors of proteins deposited in PDB as seen by a flexible protein structure alignment program FATCAT, developed previously in our group. The database, searchable by a protein PDB code, provides lists of proteins with statistically significant structural similarity and on lower menu levels provides detailed alignments, interactive superposition of structures and positions of hinges that were identified in the comparison. While superficially similar to other structural protein alignment resources, FSN provides a unique resource to study not only protein structural similarity, but also how protein structures change. FSN is available from a server http://fatcat.burnham.org/fatcat/struct_neighbor and by direct links from the PDB database.

Databases, Protein↗

TLFAM--a new set of protein family databases.

PFAM is a popular and effective database of Hidden Markov Models (HMMs), which represent a wide range of protein families. Here, we introduce TLFAM as a more specific set of HMM databases. Analyses of bacterial genomes using TLFAM-Pro show better scores, E-values, and alignment lengths than those using the more generalized PFAM. Since PFAM will still find hits that TLFAM-Pro will not, we recommend that they be used jointly, rather than exclusively. This method provides the best features of both databases. This method has been extended to a number of other organism types, such as archaea, and the databases are freely available to interested researchers.

Databases, Protein↗