PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

CEAS: cis-regulatory element annotation system.

The recent availability of high-density human genome tiling arrays enables biologists to conduct ChIP-chip experiments to locate the in vivo-binding sites of transcription factors in the human genome and explore the regulatory mechanisms. Once genomic regions enriched by transcription factor ChIP-chip are located, genome-scale downstream analyses are crucial but difficult for biologists without strong bioinformatics support. We designed and implemented the first web server to streamline the ChIP-chip downstream analyses. Given genome-scale ChIP regions, the cis-regulatory element annotation system (CEAS) retrieves repeat-masked genomic sequences, calculates GC content, plots evolutionary conservation, maps nearby genes and identifies enriched transcription factor-binding motifs. Biologists can utilize CEAS to retrieve useful information for ChIP-chip validation, assemble important knowledge to include in their publication and generate novel hypotheses (e.g. transcription factor cooperative partner) for further study. CEAS helps the adoption of ChIP-chip in mammalian systems and provides insights towards a more comprehensive understanding of transcriptional regulatory mechanisms. The URL of the server is http://ceas.cbi.pku.edu.cn.

Binding Sites↗

Genotyping and annotation of Affymetrix SNP arrays.

In this paper we develop a new method for genotyping Affymetrix single nucleotide polymorphism (SNP) array. The method is based on (i) using multiple arrays at the same time to determine the genotypes and (ii) a model that relates intensities of individual SNPs to each other. The latter point allows us to annotate SNPs that have poor performance, either because of poor experimental conditions or because for one of the alleles the probes do not behave in a dose-response manner. Generally, our method agrees well with a method developed by Affymetrix. When both methods make a call they agree in 99.25% (using standard settings) of the cases, using a sample of 113 Affymetrix 10k SNP arrays. In the majority of cases where the two methods disagree, our method makes a genotype call, whereas the method by Affymetrix makes a no call, i.e. the genotype of the SNP is not determined. By visualization it is indicated that our method is likely to be correct in majority of these cases. In addition, we demonstrate that our method produces more SNPs that are in concordance with Hardy-Weinberg equilibrium than the method by Affymetrix. Finally, we have validated our method on HapMap data and shown that the performance of our method is comparable to other methods.

Algorithms↗

Robust analysis of 5'-transcript ends (5'-RATE): a novel technique for transcriptome analysis and genome annotation.

Complicated cloning procedures and the high cost of sequencing have inhibited the wide application of serial analysis of gene expression and massively parallel signature sequencing for genome-wide transcriptome profiling of complex genomes. Here we describe a new method called robust analysis of 5'-transcript ends (5'-RATE) for rapid and cost-effective isolation of long 5' transcript ends (approximately 80 bp). It consists of three major steps including 5'-oligocapping of mRNA, NlaIII tag and ditag generation, and pyrosequencing of NlaIII tags. Complicated steps, such as purification and cloning of concatemers, colony picking and plasmid DNA purification, are eliminated and the conventional Sanger sequencing method is replaced with the newly developed pyrosequencing method. Sequence analysis of a maize 5'-RATE library revealed complex alternative transcription start sites and a 5' poly(A) tail in maize transcripts. Our results demonstrate that 5'-RATE is a simple, fast and cost-effective method for transcriptome analysis and genome annotation of complex genomes.

5' Untranslated Regions↗

Differential annotation of tRNA genes with anticodon CAT in bacterial genomes.

We have developed three strategies to discriminate among the three types of tRNA genes with anticodon CAT (tRNA(Ile), elongator tRNA(Met) and initiator tRNA(fMet)) in bacterial genomes. With these strategies, we have classified the tRNA genes from 234 bacterial and several organellar genomes. These sequences, in an aligned or unaligned format, may be used for the identification and annotation of tRNA (CAT) genes in other genomes. The first strategy is based on the position of the problem sequences in a phenogram (a tree-like network), the second on the minimum average number of differences against the tRNA sequences of the three types and the third on the search for the highest score value against the profiles of the three types of tRNA genes. The species with the maximum number of tRNA(fMet) and tRNA(Met) was Photobacterium profundum, whereas the genome of one Escherichia coli strain presented the maximum number of tRNA(Ile) (CAT) genes. This last tRNA gene and tilS, encoding an RNA-modifying enzyme, are not essential in bacteria. The acquisition of a tRNA(Ile) (TAT) gene by Mycoplasma mobile has led to the loss of both the tRNA(Ile) (CAT) and the tilS genes. The new tRNA has appropriated the function of decoding AUA codons.

Anticodon↗

MACiE (Mechanism, Annotation and Classification in Enzymes): novel tools for searching catalytic mechanisms.

MACiE (Mechanism, Annotation and Classification in Enzymes) is a database of enzyme reaction mechanisms, and is publicly available as a web-based data resource. This paper presents the first release of a web-based search tool to explore enzyme reaction mechanisms in MACiE. We also present Version 2 of MACiE, which doubles the dataset available (from Version 1). MACiE can be accessed from http://www.ebi.ac.uk/thornton-srv/databases/MACiE/

Catalysis↗

The National Microbial Pathogen Database Resource (NMPDR): a genomics platform based on subsystem annotation.

The National Microbial Pathogen Data Resource (NMPDR) (http://www.nmpdr.org) is a National Institute of Allergy and Infections Disease (NIAID)-funded Bioinformatics Resource Center that supports research in selected Category B pathogens. NMPDR contains the complete genomes of approximately 50 strains of pathogenic bacteria that are the focus of our curators, as well as >400 other genomes that provide a broad context for comparative analysis across the three phylogenetic Domains. NMPDR integrates complete, public genomes with expertly curated biological subsystems to provide the most consistent genome annotations. Subsystems are sets of functional roles related by a biologically meaningful organizing principle, which are built over large collections of genomes; they provide researchers with consistent functional assignments in a biologically structured context. Investigators can browse subsystems and reactions to develop accurate reconstructions of the metabolic networks of any sequenced organism. NMPDR provides a comprehensive bioinformatics platform, with tools and viewers for genome analysis. Results of precomputed gene clustering analyses can be retrieved in tabular or graphic format with one-click tools. NMPDR tools include Signature Genes, which finds the set of genes in common or that differentiates two groups of organisms. Essentiality data collated from genome-wide studies have been curated. Drug target identification and high-throughput, in silico, compound screening are in development.

Bacteria↗

Snap: an integrated SNP annotation platform.

Snap (Single Nucleotide Polymorphism Annotation Platform) is a server designed to comprehensively analyze single genes and relationships between genes basing on SNPs in the human genome. The aim of the platform is to facilitate the study of SNP finding and analysis within the framework of medical research. Using a user-friendly web interface, genes can be searched by name, description, position, SNP ID or clone name. Several public databases are integrated, including gene information from Ensembl, protein features from Uniprot/SWISS-PROT, Pfam and DAS-CBS. Gene relationships are fetched from BIND, MINT, KEGG and are integrated with ortholog data from TreeFam to extend the current interaction networks. Integrated tools for primer-design and mis-splicing analysis have been developed to facilitate experimental analysis of individual genes with focus on their variation. Snap is available at http://snap.humgen.au.dk/ and at http://snap.genomics.org.cn/.

Databases, Nucleic Acid↗

The SBASE domain library: a collection of annotated protein segments.

SBASE is a database of annotated protein domain sequences representing various structural, functional, ligand binding and topogenic segments of proteins. The current release of SBASE contains 27,211 entries which are provided with standardized names in order to facilitate retrieval. SBASE is cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalog of protein sequence patterns [Bairoch, A. (1992) Nucleic Acids Res., 20, Suppl., 2013-2118]. SBASE can be used to establish domain homologies through database search using programs such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], FASTDB [Brutlag et al. (1990) Comp. Appl. Biosci., 6, 237-245] or BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513], which is especially useful in the case of loosely defined domain types for which efficient consensus patterns cannot be established. The use of SBASE is illustrated on the DNA binding protein Brain-4. The database and a set of search and retrieval tools are freely available on request to the authors or by anonymous 'ftp' file transfer from < ftp.icgeb.trieste.it >.

Amino Acid Sequence↗

Annotated contract technic: special adolescent groups.

Fifty-six consecutive adolescent inpatients at the Adolescent Center of Houston International Hospital were subjects of a prospective study reviewing the efficacy of the Annotated Contract technic. Three primary diagnostic groups evolved: adjustment reaction of adolescence (ARA); borderline schizophrenia (BL-S); and antisocial personality (A-P). Follow-up of 6 to 20 months showed the technic overall to be extremely useful for ARA group, useful for BL-S group, and of questionable value for A-P group. The contract technic was initially accepted by all groups; it was facilitative in family conferences for 89%, 80%, and 56% of the ARA, BL-S, and A-P groups respectively. Most patients were able to use the contract as a barometer of their own progress; the A-P group used the contract largely to ventilate hostilities and/or limit parental demands. Only 18% of the ARA, compared to 50% of BL-S and 94% of A-P groups, ultimately ignored the contracts altogether. The technic is adaptable for clinical research as well as office practice.

Adolescent↗

Measurement of return on investment of workplace education: an annotated list of references.

In the context of downsizing and with the decline of revenue in healthcare organizations, educators recognize the need to develop strategies to measure the subsidy of staff development programs to the organization's welfare. An annotated reference list is provided to assist educators who have little time for an extensive literature review with a place to begin development of a plan to measure return on investment.

Clinical Competence↗

Natural variation among human adenoviruses: genome sequence and annotation of human adenovirus serotype 1.

The 36,001 base pair DNA sequence of human adenovirus serotype 1 (HAdV-1) has been determined, using a 'leveraged primer sequencing strategy' to generate high quality sequences economically. This annotated genome (GenBank AF534906) confirms anticipated similarity to closely related species C (formerly subgroup), human adenoviruses HAdV-2 and -5, and near identity with earlier reports of sequences representing parts of the HAdV-1 genome. A first round of HAdV-1 sequence data acquisition used PCR amplification and sequencing primers from sequences common to the genomes of HAdV-2 and -5. The subsequent rounds of sequencing used primers derived from the newly generated data. Corroborative re-sequencing with primers selected from this HAdV-1 dataset generated sparsely tiled arrays of high quality sequencing ladders spanning both complementary strands of the HAdV-1 genome. These strategies allow for rapid and accurate low-pass sequencing of genomes. Such rapid genome determinations facilitate the development of specific probes for differentiation of family, serotype, subtype and strain (e.g. pathogen genome signatures). These will be used to monitor epidemic outbreaks of acute respiratory disease in a defined test bed by the Epidemic Outbreak Surveillance (EOS) project.

Adenovirus Infections, Human↗

The Annotated Blueprint: Integrated Functional Genomic Resources for a model Tetraploid Wheat Triticum turgidum cv. Kronos.

Triticum turgidum cv. Kronos is a tetraploid wheat cultivar that underpins one of the richest community platforms for functional genomics. Over the past decade, about 3,000 exome- and promoter-capture datasets, linked to mutagenized seed stocks, and transcriptomic and phenotypic resources have accumulated, yet the absence of a reference genome has constrained their impact. Here, we present a chromosome-scale reference genome of Kronos with high-confidence annotations, including manual curation of over 1,000 disease resistance (NLR) genes. This reference revealed previously hidden NLR diversity and clarified their genomic organization at chromosomal ends. Re-analysis of exome- and promoter-capture datasets enabled high-resolution mutation discovery in genes and regulatory regions that were previously inaccessible, uncovering the full standing variation present in Kronos mutant lines. We further re-curated transcriptomic and small RNA datasets, generating improved, genome-wide maps of microRNAs and phasiRNAs important for wheat development. Collectively, these resources elevate Kronos to reference quality and establish it as a versatile platform for functional and translational wheat research.

Journal Article↗

An annotated catalog of inverted repeats of Caenorhabditis elegans chromosomes III and X, with observations concerning odd/even biases and conserved motifs.

We have taken a computational approach to the problem of discovering and deciphering the grammar and syntax of gene regulation in eukaryotes. A logical first step is to produce an annotated catalog of all regulatory sites in a given genome. Likely candidates for such sites are direct and indirect repeats, including three subcategories of indirect repeats: inverted (palindromic), everted, and mirror-image repeats. To that end we have produced a searchable database of inverted repeats of chromosomes III and X of Caenorhabditis elegans, the first completely sequenced multicellular eukaryote. Initial results from the use of this catalog are observations concerning odd/even biases in perfect IRs. The potential usefulness of the catalog as a discovery tool for promoters was shown for some of the genes involved with G-protein functions and for heat shock protein 104 (hsp104).

Algorithms↗

Transcript mapping and genome annotation of ascidian mtDNA using EST data.

Mitochondrial transcripts of two ascidian species were reconstructed through sequence assembly of publicly available ESTs resembling mitochondrial DNA sequences (mt-ESTs). This strategy allowed us to analyze processing and mapping of the mitochondrial transcripts and to investigate the gene organization of a previously uncharacterized mitochondrial genome (mtDNA). This new strategy would greatly facilitate the sequencing and annotation of mtDNAs. In Ciona intestinalis, the assembled mt-ESTs covered 22 mitochondrial genes ( approximately 12,000 bp) and provided the partial sequence of the mtDNA and the prediction of its gene organization. Such sequences were confirmed by amplification and sequencing of the entire Ciona mtDNA. For Halocynthia roretzi, for which the mtDNA sequence was already available, the inferred mt transcripts allowed better definition of gene boundaries (16S rRNA, ND1, ATP6, and tRNA-Ser genes) and the identification of a new gene (an additional Phe-tRNA). In both species, polycistronic and immature transcripts, creation of stop codons by polyadenylation, tRNA signal processing, and rRNA transcript termination signals were identified, thus suggesting that the main features of mitochondrial transcripts are conserved in Chordata.

Animals↗

Toward a functional annotation of the human genome using artificial transcription factors.

We have developed a novel, high-throughput approach to collecting randomly perturbed gene-expression profiles from the human genome.A human 293 cell library that stably expresses randomly chosen zinc-finger transcription factors was constructed, and the expression profile of each cell line was obtained using cDNA microarray technology.Gene expression profiles from a total of 132 cell lines were collected and analyzed by (1) a simple clustering method based on expression-profile similarity, and (2) the shortest-path analysis method. These analyses identified a number of gene groups, and further investigation revealed that the genes that were grouped together had close biological relationships. The artificial transcription factor-based random genome perturbation method thus provides a novel functional genomic tool for annotation and classification of genes in the human genome and those of many other organisms.

Antigens, Neoplasm↗

Analysis and functional annotation of an expressed sequence tag collection for tropical crop sugarcane.

To contribute to our understanding of the genome complexity of sugarcane, we undertook a large-scale expressed sequence tag (EST) program. More than 260,000 cDNA clones were partially sequenced from 26 standard cDNA libraries generated from different sugarcane tissues. After the processing of the sequences, 237,954 high-quality ESTs were identified. These ESTs were assembled into 43,141 putative transcripts. Of the assembled sequences, 35.6% presented no matches with existing sequences in public databases. A global analysis of the whole SUCEST data set indicated that 14,409 assembled sequences (33% of the total) contained at least one cDNA clone with a full-length insert. Annotation of the 43,141 assembled sequences associated almost 50% of the putative identified sugarcane genes with protein metabolism, cellular communication/signal transduction, bioenergetics, and stress responses. Inspection of the translated assembled sequences for conserved protein domains revealed 40,821 amino acid sequences with 1415 Pfam domains. Reassembling the consensus sequences of the 43,141 transcripts revealed a 22% redundancy in the first assembling. This indicated that possibly 33,620 unique genes had been identified and indicated that >90% of the sugarcane expressed genes were tagged.

Computational Biology↗

Generation, annotation, evolutionary analysis, and database integration of 20,000 unique sea urchin EST clusters.

Together with the hemichordates, sea urchins represent basal groups of nonchordate invertebrate deuterostomes that occupy a key position in bilaterian evolution. Because sea urchin embryos are also amenable to functional studies, the sea urchin system has emerged as one of the leading models for the analysis of the function of genomic regulatory networks that control development. We have analyzed a total of 107,283 cDNA clones of libraries that span the development of the sea urchin Strongylocentrotus purpuratus. Normalization by oligonucleotide fingerprinting, EST sequencing and sequence clustering resulted in an EST catalog comprised of 20,000 unique genes or gene fragments. Around 7000 of the unique EST consensus sequences were associated with molecular and developmental functions. Phylogenetic comparison of the identified genes to the genome of the urochordate Ciona intestinalis indicate that at least one quarter of the genes thought to be chordate specific were already present at the base of deuterostome evolution. Comparison of the number of gene copies in sea urchins to those in chordates and vertebrates indicates that the sea urchin genome has not undergone extensive gene or complete genome duplications. The established unique gene set represents an essential tool for the annotation and assembly of the forthcoming sea urchin genome sequence. All cDNA clones and filters of all analyzed libraries are available from the resource center of the German genome project at http://www.rzpd.de.

Animals↗

Visualization of multiple genome annotations and alignments with the K-BROWSER.

We introduce a novel genome browser application, the K-BROWSER, that allows intuitive visualization of biological information across an arbitrary number of multiply aligned genomes. In particular, the K-BROWSER simultaneously displays an arbitrary number of genomes both through overlaid annotations and predictions that describe their respective characteristics, and through the multiple alignment that describes their global relationship to one another. The browsing environment has been designed to allow users seamless access to information available in every genome and, furthermore, to allow easy navigation within and between genomes. As of the date of publication, the K-BROWSER has been set up on the human, mouse, and rat genomes.

Animals↗