PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Functional annotation of proteomic sequences based on consensus of sequence and structural analysis.

To maximise the assignment of function of the proteins encoded by a genome and to aid the search for novel drug targets, there is an emerging need for sensitive methods of predicting protein function on a genome-wide basis. GeneAtlas is an automated, high-throughput pipeline for the prediction of protein structure and function using sequence similarity detection, homology modelling and fold recognition methods. GeneAtlas is described in detail here. To test GeneAtlas, a 'virtual' genome was used, a subset of PDB structures from the SCOP database, in which the functional relationships are known. GeneAtlas detects additional relationships by building 3D models in comparison with the sequence searching method PSI-BLAST. Functionally related proteins with sequence identity below the twilight zone can be recognised correctly.

Consensus Sequence↗

Annotation pattern of ESTs from Spodoptera frugiperda Sf9 cells and analysis of the ribosomal protein genes reveal insect-specific features and unexpectedly low codon usage bias.

MOTIVATION: A whole set of Expressed Sequence Tags (ESTs) from the Sf9 cell line of Spodoptera frugiperda is presented here for the first time. By this way we want to identify both conserved and specific genes of this pest species. We also expect from this analysis to find a class of protein sequences providing a tool to explore genomic features and phylogeny of Lepidoptera. RESULTS: The ESTs display both housekeeping as well as developmentally regulated genes, and a high percentage of sequences with unknown function. Among the identified ORFs, almost all ribosomal proteins (RPs) were found with high EST redundancy and hence sequence accuracy. The codon usage found among RP genes is in average surprisingly much less biased in Lepidoptera than in other organisms. Other Spodoptera genes also displayed a low bias, suggesting a general genome expression feature in this Lepidoptera. We also found that the L35A and L36 RP sequences, respectively, display 40 and 10 amino-acid insertions, both being present only in insects. Sequence analysis suggests that they are probably not subjected to a strong selective pressure and may be good phylogenetic markers for Lepidoptera. Most interestingly, the Lepidoptera sequences of 9 RP genes displayed a specific signature different from the canonical one. We conclude that the RP family allows valuable comparative genomics and phylogeny of Lepidoptera. AVAILABILITY: All EST sequence data are available from the private 'Spodo-Base' upon request.

Abstracting and Indexing↗

PMUT: a web-based tool for the annotation of pathological mutations on proteins.

PMUT allows the fast and accurate prediction (approximately 80% success rate in humans) of the pathological character of single point amino acidic mutations based on the use of neural networks. The program also allows the fast scanning of mutational hot spots, which are obtained by three procedures: (1) alanine scanning, (2) massive mutation and (3) genetically accessible mutations. A graphical interface for Protein Data Bank (PDB) structures, when available, and a database containing hot spot profiles for all non-redundant PDB structures are also accessible from the PMUT server.

Amino Acid Substitution↗

MineBlast: a literature presentation service supporting protein annotation by data mining of BLAST results.

MineBlast is a web service for literature search and presentation based on data-mining results received from UniProt. Users can submit a simple list of protein sequences via a web-based interface. MineBlast performs a BLASTP search in UniProt to identify names and synonyms based on homologous proteins and subsequently queries PubMed, using combined search terms inorder to find and present relevant literature.

Database Management Systems↗

GARSA: genomic analysis resources for sequence annotation.

SUMMARY: Growth of genome data and analysis possibilities have brought new levels of difficulty for scientists to understand, integrate and deal with all this ever-increasing information. In this scenario, GARSA has been conceived aiming to facilitate the tasks of integrating, analyzing and presenting genomic information from several bioinformatics tools and genomic databases, in a flexible way. GARSA is a user-friendly web-based system designed to analyze genomic data in the context of a pipeline. EST and GGS data can be analyzed using the system since it accepts (1) chromatograms, (2) download of sequences from GenBank, (3) Fasta files stored locally or (4) a combination of all three. Quality evaluation of chromatograms, vector removing and clusterization are easily performed as part of the pipeline. A number of local and customizable Blast and CDD analyses can be performed as well as Interpro, complemented with phylogeny analyses. GARSA is being used for the analyses of Trypanosoma vivax (GSS and EST), Trypanosoma rangeli (GSS, EST and ORESTES), Bothrops jararaca (EST), Piaractus mesopotamicus (EST) and Lutzomyia longipalpis (EST). AVAILABILITY: The GARSA system is freely available under GPL license (http://www.biowebdb.org/garsa/). For download requests visit http://www.biowebdb.org/garsa/ or contact Dr Alberto Dávila.

Animals↗

Sequence-based heuristics for faster annotation of non-coding RNA families.

MOTIVATION: Non-coding RNAs (ncRNAs) are functional RNA molecules that do not code for proteins. Covariance Models (CMs) are a useful statistical tool to find new members of an ncRNA gene family in a large genome database, using both sequence and, importantly, RNA secondary structure information. Unfortunately, CM searches are extremely slow. Previously, we created rigorous filters, which provably sacrifice none of a CM's accuracy, while making searches significantly faster for virtually all ncRNA families. However, these rigorous filters make searches slower than heuristics could be. RESULTS: In this paper we introduce profile HMM-based heuristic filters. We show that their accuracy is usually superior to heuristics based on BLAST. Moreover, we compared our heuristics with those used in tRNAscan-SE, whose heuristics incorporate a significant amount of work specific to tRNAs, where our heuristics are generic to any ncRNA. Performance was roughly comparable, so we expect that our heuristics provide a high-quality solution that--unlike family-specific solutions--can scale to hundreds of ncRNA families. AVAILABILITY: The source code is available under GNU Public License at the supplementary web site.

Algorithms↗

Automated discovery of 3D motifs for protein function annotation.

MOTIVATION: Function inference from structure is facilitated by the use of patterns of residues (3D motifs), normally identified by expert knowledge, that correlate with function. As an alternative to often limited expert knowledge, we use machine-learning techniques to identify patterns of 3-10 residues that maximize function prediction. This approach allows us to test the assumption that residues that provide function are the most informative for predicting function. RESULTS: We apply our method, GASPS, to the haloacid dehalogenase, enolase, amidohydrolase and crotonase superfamilies and to the serine proteases. The motifs found by GASPS are as good at function prediction as 3D motifs based on expert knowledge. The GASPS motifs with the greatest ability to predict protein function consist mainly of known functional residues. However, several residues with no known functional role are equally predictive. For four groups, we show that the predictive power of our 3D motifs is comparable with or better than approaches that use the entire fold (Combinatorial-Extension) or sequence profiles (PSI-BLAST). AVAILABILITY: Source code is freely available for academic use by contacting the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

SChiSM2: creating interactive web page annotations of molecular structure models using Jmol.

UNLABELLED: SChiSM2 is a web server-based program for creating web pages that include interactive molecular graphics using the freely-available applet, Jmol, for illustration. The program works with Internet Explorer and Firefox on Windows, Safari and Firefox on Mac OSX and Firefox on Linux. AVAILABILITY: The program can be accessed at the following address: http://ci.vbi.vt.edu/cammer/schism2.html.

Algorithms↗