PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Functional analysis and annotation of the virulence plasmid pMUM001 from Mycobacterium ulcerans.

The presence of a 174 kb plasmid called pMUM001 in Mycobacterium ulcerans, the first example of a mycobacterial plasmid encoding a virulence determinant, was recently reported. Over half of pMUM001 is devoted to six genes, three of which encode giant polyketide synthases (PKS) that produce mycolactone, an unusual cytotoxic lipid produced by M. ulcerans. In this present study the remaining 75 non-PKS-associated protein-coding sequences (CDS) are analysed and it is shown that pMUM001 is a low-copy-number element with a functional ori that supports replication in Mycobacterium marinum but not in the fast-growing mycobacteria Mycobacterium smegmatis and Mycobacterium fortuitum. Sequence analyses revealed a highly mosaic plasmid gene structure that is reminiscent of other large plasmids. Insertion sequences (IS) and fragments of IS, some previously unreported, are interspersed among functional gene clusters, such as those genes involved in plasmid replication, the synthesis of mycolactone, and a potential phosphorelay signal transduction system. Among the IS present on pMUM001 were multiple copies of the high-copy-number M. ulcerans elements IS2404 and IS2606. No plasmid transfer systems were identified, suggesting that trans-acting factors are required for mobilization. The results presented here provide important insights into this unusual virulence plasmid from an emerging but neglected human pathogen.

Bacterial Proteins↗

Genome annotation past, present, and future: how to define an ORF at each locus.

Driven by competition, automation, and technology, the genomics community has far exceeded its ambition to sequence the human genome by 2005. By analyzing mammalian genomes, we have shed light on the history of our DNA sequence, determined that alternatively spliced RNAs and retroposed pseudogenes are incredibly abundant, and glimpsed the apparently huge number of non-coding RNAs that play significant roles in gene regulation. Ultimately, genome science is likely to provide comprehensive catalogs of these elements. However, the methods we have been using for most of the last 10 years will not yield even one complete open reading frame (ORF) for every gene--the first plateau on the long climb toward a comprehensive catalog. These strategies--sequencing randomly selected cDNA clones, aligning protein sequences identified in other organisms, sequencing more genomes, and manual curation--will have to be supplemented by large-scale amplification and sequencing of specific predicted mRNAs. The steady improvements in gene prediction that have occurred over the last 10 years have increased the efficacy of this approach and decreased its cost. In this Perspective, I review the state of gene prediction roughly 10 years ago, summarize the progress that has been made since, argue that the primary ORF identification methods we have relied on so far are inadequate, and recommend a path toward completing the Catalog of Protein Coding Genes, Version 1.0.

Animals↗

Judging the quality of gene expression-based clustering methods using gene annotation.

We compare several commonly used expression-based gene clustering algorithms using a figure of merit based on the mutual information between cluster membership and known gene attributes. By studying various publicly available expression data sets we conclude that enrichment of clusters for biological function is, in general, highest at rather low cluster numbers. As a measure of dissimilarity between the expression patterns of two genes, no method outperforms Euclidean distance for ratio-based measurements, or Pearson distance for non-ratio-based measurements at the optimal choice of cluster number. We show the self-organized-map approach to be best for both measurement types at higher numbers of clusters. Clusters of genes derived from single- and average-linkage hierarchical clustering tend to produce worse-than-random results.

Algorithms↗

Human chromosome-specific cDNA libraries: new tools for gene identification and genome annotation.

To date, only a small percentage of human genes have been cloned and mapped. To facilitate more rapid gene mapping and disease gene isolation, chromosome 5-specific cDNA libraries have been constructed from five sources. DNA sequencing and regional mapping of 205 unique cDNAs indicates that 25 are from known chromosome 5 genes and 138 are from new chromosome 5 genes (a frequency of 79.5%). Sequence complexity estimates indicate that each library contains -20% of the approximately 5000 genes that are believed to reside on chromosome 5. This study more than doubles the number of genes mapped to chromosome 5 and describes an important new tool for disease gene isolation.

Base Sequence↗

Massive sequence comparisons as a help in annotating genomic sequences.

An all-by-all comparison of all the publicly available protein sequences from plants has been performed, followed by a clusterization process. Within each of the 1064 resulting clusters-containing sequences that are orthologous as well as paralogous-the sequences have been submitted to a pyramidal classification and their domains delineated by an automated procedure à la. This process provides a means for easily checking for any apparent inconsistency in a cluster, for example, whether one sequence is shorter or longer than the others, one domain is missing, etc. In such cases, the alignment of the DNA sequence of the gene with that of a close homologous protein often reveals (in 10% of the clusters) probable sequencing errors (leading to frameshifts) or probable wrong intron/exon predictions. The composition of the clusters, their pyramidal classifications, and domain decomposition, as well as our comments when appropriate, are available from http://chlora.infobiogen.fr:1234/PHYTOPROT.

Amino Acid Sequence↗

Genome-wide annotation and expression profiling of cell cycle regulatory genes in Chlamydomonas reinhardtii.

Eukaryotic cell cycles are driven by a set of regulators that have undergone lineage-specific gene loss, duplication, or divergence in different taxa. It is not known to what extent these genomic processes contribute to differences in cell cycle regulatory programs and cell division mechanisms among different taxonomic groups. We have undertaken a genome-wide characterization of the cell cycle genes encoded by Chlamydomonas reinhardtii, a unicellular eukaryote that is part of the green algal/land plant clade. Although Chlamydomonas cells divide by a noncanonical mechanism termed multiple fission, the cell cycle regulatory proteins from Chlamydomonas are remarkably similar to those found in higher plants and metazoans, including the proteins of the RB-E2F pathway that are absent in the fungal kingdom. Unlike in higher plants and vertebrates where cell cycle regulatory genes have undergone extensive duplication, most of the cell cycle regulators in Chlamydomonas have not. The relatively small number of cell cycle genes and growing molecular genetic toolkit position Chlamydomonas to become an important model for higher plant and metazoan cell cycles.

Algal Proteins↗

Annotations and functional analyses of the rice WRKY gene superfamily reveal positive and negative regulators of abscisic acid signaling in aleurone cells.

The WRKY proteins are a superfamily of regulators that control diverse developmental and physiological processes. This family was believed to be plant specific until the recent identification of WRKY genes in nonphotosynthetic eukaryotes. We have undertaken a comprehensive computational analysis of the rice (Oryza sativa) genomic sequences and predicted the structures of 81 OsWRKY genes, 48 of which are supported by full-length cDNA sequences. Eleven OsWRKY proteins contain two conserved WRKY domains, while the rest have only one. Phylogenetic analyses of the WRKY domain sequences provide support for the hypothesis that gene duplication of single- and two-domain WRKY genes, and loss of the WRKY domain, occurred in the evolutionary history of this gene family in rice. The phylogeny deduced from the WRKY domain peptide sequences is further supported by the position and phase of the intron in the regions encoding the WRKY domains. Analyses for chromosomal distributions reveal that 26% of the predicted OsWRKY genes are located on chromosome 1. Among the dozen genes tested, OsWRKY24, -51, -71, and -72 are induced by abscisic acid (ABA) in aleurone cells. Using a transient expression system, we have demonstrated that OsWRKY24 and -45 repress ABA induction of the HVA22 promoter-beta-glucuronidase construct, while OsWRKY72 and -77 synergistically interact with ABA to activate this reporter construct. This study provides a solid base for functional genomics studies of this important superfamily of regulatory genes in monocotyledonous plants and reveals a novel function for WRKY genes, i.e. mediating plant responses to ABA.

Abscisic Acid↗

Protein surface analysis for function annotation in high-throughput structural genomics pipeline.

Structural genomics (SG) initiatives are expanding the universe of protein fold space by rapidly determining structures of proteins that were intentionally selected on the basis of low sequence similarity to proteins of known structure. Often these proteins have no associated biochemical or cellular functions. The SG success has resulted in an accelerated deposition of novel structures. In some cases the structural bioinformatics analysis applied to these novel structures has provided specific functional assignment. However, this approach has also uncovered limitations in the functional analysis of uncharacterized proteins using traditional sequence and backbone structure methodologies. A novel method, named pvSOAR (pocket and void Surface of Amino Acid Residues), of comparing the protein surfaces of geometrically defined pockets and voids was developed. pvSOAR was able to detect previously unrecognized and novel functional relationships between surface features of proteins. In this study, pvSOAR is applied to several structural genomics proteins. We examined the surfaces of YecM, BioH, and RpiB from Escherichia coli as well as the CBS domains from inosine-5'-monosphate dehydrogenase from Streptococcus pyogenes, conserved hypothetical protein Ta549 from Thermoplasm acidophilum, and CBS domain protein mt1622 from Methanobacterium thermoautotrophicum with the goal to infer information about their biochemical function.

Adenine Nucleotides↗