PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

5' Long serial analysis of gene expression (LongSAGE) and 3' LongSAGE for transcriptome characterization and genome annotation.

Complete genome annotation relies on precise identification of transcription units bounded by a transcription initiation site (TIS) and a polyadenylation site (PAS). To facilitate this process, we developed a set of two complementary methods, 5' Long serial analysis of gene expression (LS) and 3'LS. These analyses are based on the original SAGE and LS methods coupled with full-length cDNA cloning, and enable the high-throughput extraction of the first and the last 20 bp of each transcript. We demonstrate that the mapping of 5'LS and 3'LS tags to the genome allows the localization of TIS and PAS. By using 537 tag pairs mapping to the region of known genes, we confirmed that >90% of the tag pairs appropriately assigned to the first and last exons. Moreover, by using tag sequences as primers for RT-PCRs, we were able to recover putative full-length transcripts in 81% of the attempts. This large-scale generation of transcript terminal tags is at least 20-40 times more efficient than full-length cDNA cloning and sequencing in the identification of complete transcription units. The apparent precision and deep coverage makes 5'LS and 3'LS an advanced approach for genome annotation through whole-transcriptome characterization.

Animals↗

A computational and experimental approach to validating annotations and gene predictions in the Drosophila melanogaster genome.

Five years after the completion of the sequence of the Drosophila melanogaster genome, the number of protein-coding genes it contains remains a matter of debate; the number of computational gene predictions greatly exceeds the number of validated gene annotations. We have assembled a collection of >10,000 gene predictions that do not overlap existing gene annotations and have developed a process for their validation that allows us to efficiently prioritize and experimentally validate predictions from various sources by sequencing RT-PCR products to confirm gene structures. Our data provide experimental evidence for 122 protein-coding genes. Our analyses suggest that the entire collection of predictions contains only approximately 700 additional protein-coding genes. Although we cannot rule out the discovery of genes with unusual features that make them refractory to existing methods, our results suggest that the D. melanogaster genome contains approximately 14,000 protein-coding genes.

Animals↗

Annotation of cis-regulatory elements by identification, subclassification, and functional assessment of multispecies conserved sequences.

An important step toward improving the annotation of the human genome is to identify cis-acting regulatory elements from primary DNA sequence. One approach is to compare sequences from multiple, divergent species. This approach distinguishes multispecies conserved sequences (MCS) in noncoding regions from more rapidly evolving neutral DNA. Here, we have analyzed a region of approximately 238kb containing the human alpha globin cluster that was sequenced and/or annotated across the syntenic region in 22 species spanning 500 million years of evolution. Using a variety of bioinformatic approaches and correlating the results with many aspects of chromosome structure and function in this region, we were able to identify and evaluate the importance of 24 individual MCSs. This approach sensitively and accurately identified previously characterized regulatory elements but also discovered unidentified promoters, exons, splicing, and transcriptional regulatory elements. Together, these studies demonstrate an integrated approach by which to identify, subclassify, and predict the potential importance of MCSs.

Animals↗

A high-resolution annotated physical map of the human chromosome 13q12-13 region containing the breast cancer susceptibility locus BRCA2.

Various types of physical mapping data were assembled by developing a set of computer programs (Integrated Mapping Package) to derive a detailed, annotated map of a 4-Mb region of human chromosome 13 that includes the BRCA2 locus. The final assembly consists of a yeast artificial chromosome (YAC) contig with 42 members spanning the 13q12-13 region and aligned contigs of 399 cosmids established by cross-hybridization between the cosmids, which were selected from a chromosome 13-specific cosmid library using inter-Alu PCR probes from the YACs. The end sequences of 60 cosmids spaced nearly evenly across the map were used to generate sequence-tagged sites (STSs), which were mapped to the YACs by PCR. A contig framework was generated by STS content mapping, and the map was assembled on this scaffold. Additional annotation was provided by 72 expressed sequences and 10 genetic markers that were positioned on the map by hybridization to cosmids.

BRCA2 Protein↗

Annotating the human proteome.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner.

Databases, Factual↗

Hypnosis and cancer: an annotated bibliography 1985-1995.

The purpose of this annotated bibliography is to provide the reader with resources to explore the relationship between hypnosis and cancer. Items are included only if they contain explicit reference to this relationship and describe it in some detail. This bibliography includes 91 items published in English from 1985 to 1995, inclusive. For the reader's convenience, the annotations are organized into three categories: general discussions; case reports or case studies; and experimental and nonexperimental group designs.

Humans↗

Mental model construction in linear reasoning: evidence for the construction of initial annotated models.

According to the mental model theory, reasoners build an initial model representing the information given in the premises. In the context of relational reasoning, the question arises as to which kind of representation is used to cope with indeterminate or multimodel problems. The present article presents an array of possible answers arising from the initial construction of complete explicit models, partial explicit models, partial implicit models, a single "isomeric" model, or a single annotated model. Predictions generated from these views are tested in two experiments that vary the problem structure and the number of models consistent with the premises. Analyses of the premise processing times, answering times and accuracy show that the annotated model yields the best fit of the data. Implications of these findings for the mental model theory as developed for relational reasoning are discussed.

Humans↗

From annotated genomes to metabolic flux models and kinetic parameter fitting.

Significant advances in system-level modeling of cellular behavior can be achieved based on constraints derived from genomic information and on optimality hypotheses. For steady-state models of metabolic networks, mass conservation and reaction stoichiometry impose linear constraints on metabolic fluxes. Different objectives, such as maximization of growth rate or minimization of flux distance from a reference state, can be tested in different organisms and conditions. In particular, we have suggested that the metabolic properties of mutant bacterial strains are best described by an algorithm that performs a minimization of metabolic adjustment (MOMA) upon gene deletion. The increasing availability of many annotated genomes paves the way for a systematic application of these flux balance methods to a large variety of organisms. However, such a high throughput goal crucially depends on our capacity to build metabolic flux models in a fully automated fashion. Here we describe a pipeline for generating models from annotated genomes and discuss the current obstacles to full automation. In addition, we propose a framework for the integration of flux modeling results and high throughput proteomic data, which can potentially help in the inference of whole-cell kinetic parameters.

Algorithms↗

Imagene: an integrated computer environment for sequence annotation and analysis.

MOTIVATION: To be fully and efficiently exploited, data coming from sequencing projects together with specific sequence analysis tools need to be integrated within reliable data management systems. Systems designed to manage genome data and analysis tend to give a greater importance either to the data storage or to the methodological aspect, but lack a complete integration of both components. RESULTS: This paper presents a co-operative computer environment (called Imagenetrade mark) dedicated to genomic sequence analysis and annotation. Imagene has been developed by using an object-based model. Thanks to this representation, the user can directly manipulate familiar data objects through icons or lists. Imagene also incorporates a solving engine in order to manage analysis tasks. A global task is solved by successive divisions into smaller sub-tasks. During program execution, these sub-tasks are graphically displayed to the user and may be further re-started at any point after task completion. In this sense, Imagene is more transparent to the user than a traditional menu-driven package. Imagene also provides a user interface to display, on the same screen, the results produced by several tasks, together with the capability to annotate these results easily. In its current form, Imagene has been designed particularly for use in microbial sequencing projects. AVAILABILITY: Imagene best runs on SGI (Irix 6.3 or higher) workstations. It is distributed free of charge on a CD-ROM, but requires some Ilog licensed software to run. Some modules also require separate license agreements. Please contact the authors for specific academic conditions and other Unix platforms. CONTACT: imagene home page: http://wwwabi.snv.jussieu.fr/imagene

Bacillus subtilis↗

EDITtoTrEMBL: a distributed approach to high-quality automated protein sequence annotation.

SUMMARY: Many databases in molecular biology face the problem that the ever increasing rate of data production can no longer be handled by traditional methods, especially human curation. Therefore, a number of projects are currently investigating methods for automated sequence annotation. This paper describes the EBI's approach to this problem for protein sequences by integration of arbitrary analysis programs into a distributed and highly flexible environment. Our software framework allows an individual treatment of sequences depending on their particular properties, which is achieved through a high-level description of the preconditions and capabilities of analysing modules. This not only improves the overall performance of the annotation process, as unnecessary steps are avoided, but also enhances its quality since dependencies between different modules are taken into account. We have implemented a prototype and use it in the production of TrEMBL releases. AVAILABILITY: Upon request.

Algorithms↗

Artemis: sequence visualization and annotation.

SUMMARY: Artemis is a DNA sequence visualization and annotation tool that allows the results of any analysis or sets of analyses to be viewed in the context of the sequence and its six-frame translation. Artemis is especially useful in analysing the compact genomes of bacteria, archaea and lower eukaryotes, and will cope with sequences of any size from small genes to whole genomes. It is implemented in Java, and can be run on any suitable platform. Sequences and annotation can be read and written directly in EMBL, GenBank and GFF format. AVAILABITLTY: Artemis is available under the GNU General Public License from http://www.sanger.ac.uk/Software/Artemis

Databases, Factual↗

gff2ps: visualizing genomic annotations.

gff2psis a program for visualizing annotations of genomic sequences. The program takes the annotated features on a genomic sequence in GFF format as input, and produces a visual output in PostScript. While it can be used in a very simple way, it also allows for a great degree of customization through a number of options and/or customization files.

Computational Biology↗

Genquire: genome annotation browser/editor.

UNLABELLED: We present a software package, Genquire, that allows visualization, querying, hand editing, and de novo markup of complete or partially annotated genomes. The system is written in Perl/Tk and uses, where possible, existing BioPerl data models and methods for representation and manipulation of the sequence and annotation objects. An adaptor API is provided to allow Genquire to display a wide range of databases and flat files, and a plugins API provides an interface to other sequence analysis software. AVAILABILITY: Genquire v3.03 is open-source software. The code is available for download and/or contribution at http://www.bioinformatics.org/Genquire

Chromosome Mapping↗

An extensible application for assembling annotation for genomic data.

SUMMARY: AnnBuilder is an R package for assembling genomic annotation data. The system currently provides parsers to process annotation data from LocusLink, Gene Ontology Consortium, and Human Gene Project and can be extended to new data sources via user defined parsers. AnnBuilder differs from other existing systems in that it provides users with unlimited ability to assemble data from user selected sources. The products of AnnBuilder are files in XML format that can be easily used by different systems. AVAILABILITY: (http://www.bioconductor.org). Open source.

Database Management Systems↗

The Annotator's Assistant: an expert system for direct submission of genetic sequence data.

As DNA sequencing technology improves and more rapid techniques become routine in molecular biology labs, researchers need to expedite the incorporation of information into genetic sequence databases, such as GenBank, by directly submitting sequence data. The Annotator's Assistant is an expert system that runs on an IBM PC and helps the molecular biologist, who may have little knowledge of the structure or content of a GenBank entry, to construct a complete and valid sequence submission file. This expert system uses a simple molecular biology knowledge base and a selection of customized screen entry forms to guide the user through the entry and annotation of a sequence and its biological features. The system compiles information about the contributor, journal references, physical and functional characteristics of the nucleic acid, source organism and features, and checks it to eliminate incomplete answers and simple errors. Users supply input by answering direct and multiple-choice questions, selecting menu items and completing entry forms; on-line help is available. Users may also enter new or unusual information using generic forms. Several modules of the expert system were converted into Prolog programs and compiled, decreasing the running time significantly. The expert system rules and the data entry forms are easy to modify, update and customize for specific sequence classes.

Base Sequence↗

Protein family annotation in a multiple alignment viewer.

SUMMARY: The Pfaat protein family alignment annotation tool is a Java-based multiple sequence alignment editor and viewer designed for protein family analysis. The application merges display features such as dendrograms, secondary and tertiary protein structure with SRS retrieval, subgroup comparison, and extensive user-annotation capabilities. AVAILABILITY: The program and source code are freely available from the authors under the GNU General Public License at http://www.pfizerdtc.com

Amino Acid Sequence↗

ASAP: automated sequence annotation pipeline for web-based updating of sequence information with a local dynamic database.

The automated sequence annotation pipeline (ASAP) is designed to ease routine investigation of new functional annotations on unknown sequences, such as expressed sequence tags (ESTs), through querying of web-accessible resources and maintenance of a local database. The system allows easy use of the output from one search as the input for a new search, as well as the filtering of results. The database is used to store formats and parameters and information for parsing data from web sites. The database permits easy updating of format information should a site modify the format of a query or of a returned web page.

Database Management Systems↗

GENIA corpus--semantically annotated corpus for bio-textmining.

MOTIVATION: Natural language processing (NLP) methods are regarded as being useful to raise the potential of text mining from biological literature. The lack of an extensively annotated corpus of this literature, however, causes a major bottleneck for applying NLP techniques. GENIA corpus is being developed to provide reference materials to let NLP techniques work for bio-textmining. RESULTS: GENIA corpus version 3.0 consisting of 2000 MEDLINE abstracts has been released with more than 400,000 words and almost 100,000 annotations for biological terms.

Abstracting and Indexing↗