PubMed Health⌕ Search

Biomedical subjects

Anton J Enright

Publications and source records attributed to Anton J Enright.

12 recordsLinked to original sources

MagicMatch--cross-referencing sequence identifiers across databases.

MOTIVATION: At present, mapping of sequence identifiers across databases is a daunting, time-consuming and computationally expensive process, usually achieved by sequence similarity searches with strict threshold values. SUMMARY: We present a rapid and efficient method to map sequence identifiers across databases. The method uses the MD5 checksum algorithm for message integrity to generate sequence fingerprints and uses these fingerprints as hash strings to map sequences across databases. The program, called MagicMatch, is able to cross-link any of the major sequence databases within a few seconds on a modest desktop computer.

Algorithms↗

MicroRNAs regulate brain morphogenesis in zebrafish.

MicroRNAs (miRNAs) are small RNAs that regulate gene expression posttranscriptionally. To block all miRNA formation in zebrafish, we generated maternal-zygotic dicer (MZdicer) mutants that disrupt the Dicer ribonuclease III and double-stranded RNA-binding domains. Mutant embryos do not process precursor miRNAs into mature miRNAs, but injection of preprocessed miRNAs restores gene silencing, indicating that the disrupted domains are dispensable for later steps in silencing. MZdicer mutants undergo axis formation and differentiate multiple cell types but display abnormal morphogenesis during gastrulation, brain formation, somitogenesis, and heart development. Injection of miR-430 miRNAs rescues the brain defects in MZdicer mutants, revealing essential roles for miRNAs during morphogenesis.

Animals↗

BioLayout(Java): versatile network visualisation of structural and functional relationships.

Visualisation of biological networks is becoming a common task for the analysis of high-throughput data. These networks correspond to a wide variety of biological relationships, such as sequence similarity, metabolic pathways, gene regulatory cascades and protein interactions. We present a general approach for the representation and analysis of networks of variable type, size and complexity. The application is based on the original BioLayout program (C-language implementation of the Fruchterman-Rheingold layout algorithm), entirely re-written in Java to guarantee portability across platforms. BioLayout(Java) provides broader functionality, various analysis techniques, extensions for better visualisation and a new user interface. Examples of analysis of biological networks using BioLayout(Java) are presented.

Computer Graphics↗

Human MicroRNA targets.

MicroRNAs (miRNAs) interact with target mRNAs at specific sites to induce cleavage of the message or inhibit translation. The specific function of most mammalian miRNAs is unknown. We have predicted target sites on the 3' untranslated regions of human gene transcripts for all currently known 218 mammalian miRNAs to facilitate focused experiments. We report about 2,000 human genes with miRNA target sites conserved in mammals and about 250 human genes conserved as targets between mammals and fish. The prediction algorithm optimizes sequence complementarity using position-specific rules and relies on strict requirements of interspecies conservation. Experimental support for the validity of the method comes from known targets and from strong enrichment of predicted targets in mRNAs associated with the fragile X mental retardation protein in mammals. This is consistent with the hypothesis that miRNAs act as sequence-specific adaptors in the interaction of ribonuclear particles with translationally regulated messages. Overrepresented groups of targets include mRNAs coding for transcription factors, components of the miRNA machinery, and other proteins involved in translational regulation, as well as components of the ubiquitin machinery, representing novel feedback loops in gene regulation. Detailed information about target genes, target processes, and open-source software for target prediction (miRanda) is available at http://www.microrna.org. Our analysis suggests that miRNA genes, which are about 1% of all human genes, regulate protein production for 10% or more of all human genes.

3' Untranslated Regions↗

Identification of virus-encoded microRNAs.

RNA silencing processes are guided by small RNAs that are derived from double-stranded RNA. To probe for function of RNA silencing during infection of human cells by a DNA virus, we recorded the small RNA profile of cells infected by Epstein-Barr virus (EBV). We show that EBV expresses several microRNA (miRNA) genes. Given that miRNAs function in RNA silencing pathways either by targeting messenger RNAs for degradation or by repressing translation, we identified viral regulators of host and/or viral gene expression.

Animals↗

Detection of functional modules from protein interaction networks.

Complex cellular processes are modular and are accomplished by the concerted action of functional modules (Ravasz et al., Science 2002;297:1551-1555; Hartwell et al., Nature 1999;402:C47-52). These modules encompass groups of genes or proteins involved in common elementary biological functions. One important and largely unsolved goal of functional genomics is the identification of functional modules from genomewide information, such as transcription profiles or protein interactions. To cope with the ever-increasing volume and complexity of protein interaction data (Bader et al., Nucleic Acids Res 2001;29:242-245; Xenarios et al., Nucleic Acids Res 2002;30:303-305), new automated approaches for pattern discovery in these densely connected interaction networks are required (Ravasz et al., Science 2002;297:1551-1555; Bader and Hogue, Nat Biotechnol 2002;20:991-997; Snel et al., Proc Natl Acad Sci USA 2002;99:5890-5895). In this study, we successfully isolate 1046 functional modules from the known protein interaction network of Saccharomyces cerevisiae involving 8046 individual pair-wise interactions by using an entirely automated and unsupervised graph clustering algorithm. This systems biology approach is able to detect many well-known protein complexes or biological processes, without reference to any additional information. We use an extensive statistical validation procedure to establish the biological significance of the detected modules and explore this complex, hierarchical network of modular interactions from which pathways can be inferred.

Algorithms↗

MicroRNA targets in Drosophila.

BACKGROUND: The recent discoveries of microRNA (miRNA) genes and characterization of the first few target genes regulated by miRNAs in Caenorhabditis elegans and Drosophila melanogaster have set the stage for elucidation of a novel network of regulatory control. We present a computational method for whole-genome prediction of miRNA target genes. The method is validated using known examples. For each miRNA, target genes are selected on the basis of three properties: sequence complementarity using a position-weighted local alignment algorithm, free energies of RNA-RNA duplexes, and conservation of target sites in related genomes. Application to the D. melanogaster, Drosophila pseudoobscura and Anopheles gambiae genomes identifies several hundred target genes potentially regulated by one or more known miRNAs. RESULTS: These potential targets are rich in genes that are expressed at specific developmental stages and that are involved in cell fate specification, morphogenesis and the coordination of developmental processes, as well as genes that are active in the mature nervous system. High-ranking target genes are enriched in transcription factors two-fold and include genes already known to be under translational regulation. Our results reaffirm the thesis that miRNAs have an important role in establishing the complex spatial and temporal patterns of gene activity necessary for the orderly progression of development and suggest additional roles in the function of the mature organism. In addition the results point the way to directed experiments to determine miRNA functions. CONCLUSIONS: The emerging combinatorics of miRNA target sites in the 3' untranslated regions of messenger RNAs are reminiscent of transcriptional regulation in promoter regions of DNA, with both one-to-many and many-to-one relationships between regulator and target. Typically, more than one miRNA regulates one message, indicative of cooperative translational control. Conversely, one miRNA may have several target genes, reflecting target multiplicity. As a guide to focused experiments, we provide detailed online information about likely target genes and binding sites in their untranslated regions, organized by miRNA or by gene and ranked by likelihood of match. The target prediction algorithm is freely available and can be applied to whole genome sequences using identified miRNA sequences.

3' Untranslated Regions↗

Protein families and TRIBES in genome sequence space.

Accurate detection of protein families allows assignment of protein function and the analysis of functional diversity in complete genomes. Recently, we presented a novel algorithm called TribeMCL for the detection of protein families that is both accurate and efficient. This method allows family analysis to be carried out on a very large scale. Using TribeMCL, we have generated a resource called TRIBES that contains protein family information, comprising annotations, protein sequence alignments and phylogenetic distributions describing 311 257 proteins from 83 completely sequenced genomes. The analysis of at least 60 934 detected protein families reveals that, with the essential families excluded, paralogy levels are similar between prokaryotes, irrespective of genome size. The number of essential families is estimated to be between 366 and 426. We also show that the currently known space of protein families is scale free and discuss the implications of this distribution. In addition, we show that smaller families are often formed by shorter proteins and discuss the reasons for this intriguing pattern. Finally, we analyse the functional diversity of protein families in entire genome sequences. The TRIBES protein family resource is accessible at http://www.ebi.ac.uk/research/cgg/tribes/.

Algorithms↗

COmplete GENome Tracking (COGENT): a flexible data environment for computational genomics.

SUMMARY: We present a database of fully sequenced and published genomes to facilitate the re-distribution of data and ensure reproducibility of results in the field of computational genomics. For its design we have implemented an extremely simple yet powerful schema to allow linking of genome sequence data to other resources. AVAILABILITY: http://maine.ebi.ac.uk:8000/services/cogent/

Computational Biology↗

Evaluation of annotation strategies using an entire genome sequence.

MOTIVATION: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. RESULTS: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences.

Amino Acid Sequence↗

Myriads of protein families, and still counting.

From the historical record of genome sequencing, we show that the rate of discovery of new families has remained constant over time, indicating that our knowledge of sequence space is far from complete.

Animals↗

Classification schemes for protein structure and function.

We examine the structural and functional classifications of the protein universe, providing an overview of the existing classification schemes, their features and inter-relationships. We argue that a unified scheme should be based on a natural classification approach and that more comparative analyses of the present schemes are required both to understand their limitations and to help delimit the number of known protein folds and their corresponding functional roles in cells.

Animals↗