PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

BioEditor-simplifying macromolecular structure annotation.

SUMMARY: BioEditor is an application to enable scientists and educators to prepare and present structure annotations containing formatted text, graphics, sequence data, and interactive molecular views. It is intended to bridge the gap between printed journal articles and Internet presentation formats. BioEditor is relevant in the era of structural genomics, where annotation and publication could become the rate determining step in structure determination. AVAILABILITY: BioEditor is available at http://bioeditor.sdsc.edu. The Web site includes the latest version of the software for Microsoft Windows, including documentation, the opportunity to submit bug reports and suggestions, example documentaries prepared with BioEditor and a repository where users can submit documentaries for posting to the site.

Biopolymers↗

MUTAGEN: multi-user tool for annotating genomes.

SUMMARY: MUTAGEN is a free prokaryotic annotation system. It offers the advantages of genome comparison, graphical sequence browsers, search facilities and open-source for user-specific adjustments. The web-interface allows several users to access the system from standard desktop computers. The Sulfolobus acidocaldarius genome, and several plasmids and viruses have so far been analysed and annotated using MUTAGEN. AVAILABILITY: MUTAGEN is released as open-source software under GPL. The code is available for download and/or contribution at http://dac.molbio.ku.dk/bioinformatics/MUTAGEN/

Database Management Systems↗

RAD and the RAD Study-Annotator: an approach to collection, organization and exchange of all relevant information for high-throughput gene expression studies.

MOTIVATION: Gene expression array technology has become increasingly widespread among researchers who recognize its numerous promises. At the same time, bench biologists and bioinformaticians have come to appreciate increasingly the importance of establishing a collaborative dialog from the onset of a study and of collecting and exchanging detailed information on the many experimental and computational procedures using a structured mechanism. This is crucial for adequate analyses of this kind of data. RESULTS: The RNA Abundance Database (RAD; http://www.cbil.upenn.edu/RAD) provides a comprehensive MIAME-supportive infrastructure for gene expression data management and makes extensive use of ontologies. Specific details on protocols, biomaterials, study designs, etc. are collected through a user-friendly suite of web annotation forms. Software has been developed to generate MAGE-ML documents to enable easy export of studies stored in RAD to any other database accepting data in this format (e.g. ArrayExpress). RAD is part of a more general Genomics Unified Schema (http://www.gusdb.org), which includes a richly annotated gene index (http://www.allgenes.org), thus providing a platform that integrates genomic and transcriptomic data from multiple organisms. This infrastructure enables a large variety of queries that incorporate visualization and analysis tools and have been tailored to serve the specific needs of projects focusing on particular organisms or biological systems.

Abstracting and Indexing↗

Database model and specification of GermOnline Release 2.0, a cross-species community annotation knowledgebase on germ cell differentiation.

UNLABELLED: GermOnline is a web-accessible relational database that enables life scientists to make a significant and sustained contribution to the annotation of genes relevant for the fields of mitosis, meiosis, germ line development and gametogenesis across species. This novel approach to genome annotation includes a platform for knowledge submission and curation as well as microarray data storage and visualization hosted by a global network of servers. AVAILABILITY: The database is accessible at http://www.germonline.org/. For convenient world-wide access we have set up a network of servers in Europe (http://germonline.unibas.ch/; http://germonline.igh.cnrs.fr/), Japan (http://germonline.biochem.s.u-tokyo.ac.jp/) and USA (http://germonline.yeastgenome.org/). SUPPLEMENTARY INFORMATION: Extended documentation of the database is available through the link 'About GermOnline' at the websites.

Animals↗

POLYVIEW: a flexible visualization tool for structural and functional annotations of proteins.

UNLABELLED: The POLYVIEW visualization server can be used to generate protein sequence annotations, including secondary structures, relative solvent accessibilities, functional motifs and polymorphic sites. Two-dimensional graphical representations in a customizable format may be generated for both known protein structures and predictions obtained using protein structure prediction servers. POLYVIEW may be used for automated generation of pictures with structural and functional annotations for publications and proteomic on-line resources. AVAILABILITY: http://polyview.cchmc.org.

Algorithms↗

A System for Automated Bacterial (genome) Integrated Annotation--SABIA.

UNLABELLED: A web-based software suite, SABIA (System for Automated Bacterial Integrated Annotation), is described that provides a comprehensive computational support for the assembly and annotation of whole bacterial genomes from the data derived from sequencing projects. AVAILABILITY: Both SABIA and supplementary materials are available at http://www.sabia.lncc.br

Algorithms↗

ColorHOR--novel graphical algorithm for fast scan of alpha satellite higher-order repeats and HOR annotation for GenBank sequence of human genome.

MOTIVATION: GenBank data are at present lacking alpha satellite higher-order repeat (HOR) annotation. Furthermore, exact HOR consensus lengths have not been reported so far. Given the fast growth of sequence databases in the centromeric region, it is of increasing interest to have efficient tools for computational identification and analysis of HORs from known sequences. RESULTS: We develop a graphical user interface method, ColorHOR, for fast computational identification of HORs in a given genomic sequence, without requiring a priori information on the composition of the genomic sequence. ColorHOR is based on an extension of the key-string algorithm and provides a color representation of the order and orientation of HORs. For the key string, we use a robust 6 bp string from a consensus alpha satellite and its representative nature is tested. ColorHOR algorithm provides a direct visual identification of HORs (direct and/or reverse complement). In more detail, we first illustrate the ColorHOR results for human chromosome 1. Using ColorHOR we determine for the first time the HOR annotation of the GenBank sequence of the whole human genome. In addition to some HORs, corresponding to those determined previously biochemically, we find new HORs in chromosomes 4, 8, 9, 10, 11 and 19. For the first time, we determine exact consensus lengths of HORs in 10 chromosomes. We propose that the HOR assignment obtained by using ColorHOR be included into the GenBank database.

Algorithms↗

Drosophila DNase I footprint database: a systematic genome annotation of transcription factor binding sites in the fruitfly, Drosophila melanogaster.

UNLABELLED: Despite increasing numbers of computational tools developed to predict cis-regulatory sequences, the availability of high-quality datasets of transcription factor binding sites limits advances in the bioinformatics of gene regulation. Here we present such a dataset based on a systematic literature curation and genome annotation of DNase I footprints for the fruitfly, Drosophila melanogaster. Using the experimental results of 201 primary references, we annotated 1367 binding sites from 87 transcription factors and 101 target genes in the D.melanogaster genome sequence. These data will provide a rich resource for future bioinformatics analyses of transcriptional regulation in Drosophila such as constructing motif models, training cis-regulatory module detectors, benchmarking alignment tools and continued text mining of the extensive literature on transcriptional regulation in this important model organism. AVAILABILITY: http://www.flyreg.org/ CONTACT: cbergman@gen.cam.ac.uk.

Binding Sites↗

GOAnno: GO annotation based on multiple alignment.

UNLABELLED: GOAnno is a web tool that automatically annotates proteins according to the Gene Ontology (GO) using evolutionary information available in hierarchized multiple alignments. GO terms present in the aligned functional subfamily can be cross-validated and propagated to obtain highly reliable predicted GO annotation based on the GOAnno algorithm. AVAILABILITY: The web tool and a reduced version for local installation are freely available at http://igbmc.u-strasbg.fr/GOAnno/GOAnno.html SUPPLEMENTARY INFORMATION: The website supplies a detailed explanation and illustration of the algorithm at http://igbmc.u-strasbg.fr/GOAnno/GOAnnoHelp.html.

Algorithms↗

Concept-based annotation of enzyme classes.

MOTIVATION: Given the explosive growth of biomedical data as well as the literature describing results and findings, it is getting increasingly difficult to keep up to date with new information. Keeping databases synchronized with current knowledge is a time-consuming and expensive task-one which can be alleviated by automatically gathering findings from the literature using linguistic approaches. We describe a method to automatically annotate enzyme classes with disease-related information extracted from the biomedical literature for inclusion in such a database. RESULTS: Enzyme names for the 3901 enzyme classes in the BRENDA database, a repository for quantitative and qualitative enzyme information, were identified in more than 100,000 abstracts retrieved from the PubMed literature database. Phrases in the abstracts were assigned to concepts from the Unified Medical Language System (UMLS) utilizing the MetaMap program, allowing for the identification of disease-related concepts by their semantic fields in the UMLS ontology. Assignments between enzyme classes and diseases were created based on their co-occurrence within a single sentence. False positives could be removed by a variety of filters including minimum number of co-occurrences, removal of sentences containing a negation and the classification of sentences based on their semantic fields by a Support Vector Machine. Verification of the assignments with a manually annotated set of 1500 sentences yielded favorable results of 92% precision at 50% recall, sufficient for inclusion in a high-quality database. AVAILABILITY: Source code is available from the author upon request. SUPPLEMENTARY INFORMATION: ftp.uni-koeln.de/institute/biochemie/pub/brenda/info/diseaseSupp.pdf.

Algorithms↗

Literature mining and database annotation of protein phosphorylation using a rule-based system.

MOTIVATION: A large volume of experimental data on protein phosphorylation is buried in the fast-growing PubMed literature. While of great value, such information is limited in databases owing to the laborious process of literature-based curation. Computational literature mining holds promise to facilitate database curation. RESULTS: A rule-based system, RLIMS-P (Rule-based LIterature Mining System for Protein Phosphorylation), was used to extract protein phosphorylation information from MEDLINE abstracts. An annotation-tagged literature corpus developed at PIR was used to evaluate the system for finding phosphorylation papers and extracting phosphorylation objects (kinases, substrates and sites) from abstracts. RLIMS-P achieved a precision and recall of 91.4 and 96.4% for paper retrieval, and of 97.9 and 88.0% for extraction of substrates and sites. Coupling the high recall for paper retrieval and high precision for information extraction, RLIMS-P facilitates literature mining and database annotation of protein phosphorylation.

Abstracting and Indexing↗

Modularized learning of genetic interaction networks from biological annotations and mRNA expression data.

MOTIVATION: Inferring the genetic interaction mechanism using Bayesian networks has recently drawn increasing attention due to its well-established theoretical foundation and statistical robustness. However, the relative insufficiency of experiments with respect to the number of genes leads to many false positive inferences. RESULTS: We propose a novel method to infer genetic networks by alleviating the shortage of available mRNA expression data with prior knowledge. We call the proposed method 'modularized network learning' (MONET). Firstly, the proposed method divides a whole gene set to overlapped modules considering biological annotations and expression data together. Secondly, it infers a Bayesian network for each module, and integrates the learned subnetworks to a global network. An algorithm that measures a similarity between genes based on hierarchy, specificity and multiplicity of biological annotations is presented. The proposed method draws a global picture of inter-module relationships as well as a detailed look of intra-module interactions. We applied the proposed method to analyze Saccharomyces cerevisiae stress data, and found several hypotheses to suggest putative functions of unclassified genes. We also compared the proposed method with a whole-set-based approach and two expression-based clustering approaches.

Algorithms↗

DIG--a system for gene annotation and functional discovery.

SUMMARY: We describe a database and information discovery system named DIG (Duke Integrated Genomics) designed to facilitate the process of gene annotation and the discovery of functional context. The DIG system collects and organizes gene annotation and functional information, and includes tools that support an understanding of genes in a functional context by providing a framework for integrating and visualizing gene expression, protein interaction and literature-based interaction networks.

Chromosome Mapping↗

TRAP: automated classification, quantification and annotation of tandemly repeated sequences.

TRAP, the Tandem Repeats Analysis Program, is a Perl program that provides a unified set of analyses for the selection, classification, quantification and automated annotation of tandemly repeated sequences. TRAP uses the results of the Tandem Repeats Finder program to perform a global analysis of the satellite content of DNA sequences, permitting researchers to easily assess the tandem repeat content for both individual sequences and whole genomes. The results can be generated in convenient formats such as HTML and comma-separated values. TRAP can also be used to automatically generate annotation data in the format of feature table and GFF files.

Algorithms↗

The evolutionary analysis of "orphans" from the Drosophila genome identifies rapidly diverging and incorrectly annotated genes.

In genome projects of eukaryotic model organisms, a large number of novel genes of unknown function and evolutionary history ("orphans") are being identified. Since many orphans have no known homologs in distant species, it is unclear whether they are restricted to certain taxa or evolve rapidly, either because of a lack of constraints or positive Darwinian selection. Here we use three criteria for the selection of putatively rapidly evolving genes from a single sequence of Drosophila melanogaster. Thirteen candidate genes were chosen from the Adh region on the second chromosome and 1 from the tip of the X chromosome. We succeeded in obtaining sequence from 6 of these in the closely related species D. simulans and D. yakuba. Only 1 of the 6 genes showed a large number of amino acid replacements and in-frame insertions/deletions. A population survey of this gene suggests that its rapid evolution is due to the fixation of many neutral or nearly neutral mutations. Two other genes showed "normal" levels of divergence between species. Four genes had insertions/deletions that destroy the putative reading frame within exons, suggesting that these exons have been incorrectly annotated. The evolutionary analysis of orphan genes in closely related species is useful for the identification of both rapidly evolving and incorrectly annotated genes.

Animals↗

iProClass: an integrated, comprehensive and annotated protein classification database.

The iProClass database is an integrated resource that provides comprehensive family relationships and structural and functional features of proteins, with rich links to various databases. It is extended from ProClass, a protein family database that integrates PIR superfamilies and PROSITE motifs. The iProClass currently consists of more than 200,000 non-redundant PIR and SWISS-PROT proteins organized with more than 28,000 superfamilies, 2600 domains, 1300 motifs, 280 post-translational modification sites and links to more than 30 databases of protein families, structures, functions, genes, genomes, literature and taxonomy. Protein and family summary reports provide rich annotations, including membership information with length, taxonomy and keyword statistics, full family relationships, comprehensive enzyme and PDB cross-references and graphical feature display. The database facilitates classification-driven annotation for protein sequence databases and complete genomes, and supports structural and functional genomic research. The iProClass is implemented in Oracle 8i object-relational system and available for sequence search and report retrieval at http://pir.georgetown.edu/iproclass/.

Databases, Factual↗

MODBASE, a database of annotated comparative protein structure models.

MODBASE (http://guitar.rockefeller.edu/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on PSI-BLAST, IMPALA and MODELLER. MODBASE uses the MySQL relational database management system for flexible and efficient querying, and the MODVIEW Netscape plugin for viewing and manipulating multiple sequences and structures. It is updated regularly to reflect the growth of the protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different datasets. The largest dataset contains models for domains in 304 517 out of 539 171 unique protein sequences in the complete TrEMBL database (23 March 2001); only models based on significant alignments (PSI-BLAST E-value < 10(-4)) and models assessed to have the correct fold are included. Other datasets include models for target selection and structure-based annotation by the New York Structural Genomics Research Consortium, models for prediction of genes in the Drosophila melanogaster genome, models for structure determination of several ribosomal particles and models calculated by the MODWEB comparative modeling web server.

Animals↗

rSNP_Guide, a database system for analysis of transcription factor binding to DNA with variations: application to genome annotation.

The analysis of gene regulatory networks has become one of the most challenging problems of the postgenomic era. Earlier we developed rSNP_Guide (http://util.bionet.nsc.ru/databases/rsnp.html), a computer system and database devoted to prediction of transcription factor (TF) binding sites (TF sites), which can be responsible for disease phenotypes. The prediction results were confirmed by 70 known relationships between TF sites and diseases, as well as by site-directed mutagenesis data. The rSNP_Guide is being investigated as a tool for TF site annotation. Previously analyzed and characterized cases of altered TF sites were used to annotate potential sites of the same type and at the same location in homologous genes. Based on 20 TF sites with known alterations in TF binding to DNA, we localized 245 potential TF sites in homologous genes. For these potential TF sites, rSNP_Guide estimates TF-DNA interaction according to three categories: 'present', 'weak', and 'absent'. The significance of each assignment is statistically measured.

Binding Sites↗