PubMed Health⌕ Search

PubMed · 16306393

Mining sequence annotation databanks for association patterns.

Abstract

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Irena I Artamonova, Goar Frishman, Mikhail S Gelfand, Dmitrij Frishman. 2005-11-01. Mining sequence annotation databanks for association patterns.. https://doi.org/10.1093/bioinformatics%2Fbti1206

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Base-pair resolution conservation data improves cell type specific sequence-to-expression prediction.

MOTIVATION: Genomic sequence-to-activity models can decipher gene regulatory mechanisms and predict the functional impact of regulatory variants. However, current models struggle to integrate information from sequences outside promoters, especially information from cell type specific regulatory elements. RESULTS: Here, we propose incorporating base-pair resolution evolutionary conservation data into genomic sequence-to-expression predictors. We explore two training strategies-training from scratch or fine-tuning an existing sequence-only model with additional conservation input. We find that in both cases, base-pair resolution conservation data improves cell type specific sequence-to-expression prediction, with training from scratch yielding the greatest benefit. The improvement in cell type specific expression prediction can be attributed in part to the fact that models trained on sequence and conservation data learn to better recognize cell type specific regulatory elements than models trained on sequence alone. AVAILABILITY: Code is available at https://github.com/ni-lab/basenji-phyloP.

Conserved Sequence↗

Lift&Add-rapid and robust addition of new species to alignments of conserved non-coding sequences.

MOTIVATION: Identifying sequence constraint across long evolutionary distances is a powerful method for the discovery of functional genomic sequences, especially putative non-coding elements. Conserved elements have been a mainstay of comparative genomic research, and can be further investigated for species-specific sequence acceleration to dissect the genetic basis of trait evolution. The conclusions of these comparative genomic studies are contingent on the number and range of species included in this phylogenetic analysis. However, while the number of metazoan genomes sequences is increasing rapidly, adding new genomes to existing whole-genome alignments remains computationally expensive. RESULTS: Here, we present a bioinformatic workflow, Lift&Add, that enables conserved elements, coding or non-coding, to be rapidly mapped to new genomes ("Lift") and subsequently be added to pre-existing multiple species alignments ("Add"), thus providing an avenue for easy exploration of these putative functional elements. Focusing here on a group of species that has been largely under-represented in genomic comparisons, the marsupials, we demonstrate the intuition behind this workflow and provide an example comparative genomic analysis that can be performed. IMPLEMENTATION AND AVAILABILITY: Lift&Add is implemented as a series of scripts in Snakemake and bash, which can be downloaded from https://github.com/navyashukladr/Lift_and_Add.

Conserved Sequence↗

SNF1/AMPK/SnRK1 kinases, global regulators at the heart of energy control?

The SNF1-related kinases are considered to be crucial elements of transcriptional, metabolic and developmental regulation in response to stress. In yeast, SNF1 is one of the main regulators in the shift from fermentation to aerobic metabolism; AMPK, its mammalian counterpart, is a master metabolic regulator involved in a variety of metabolic disorders such as diabetes and obesity. The aim of this review is to examine the literature concerning SnRK1 proteins, the SNF1 homologues in plants. The remarkable structural similarities between the plant complexes and those of yeast and mammalian suggest the existence of a common ancestral function in the regulation of energy and carbon metabolism. We will also highlight some distinctive features acquired by the plant proteins during evolution.

Conserved Sequence↗