PubMed Health⌕ Search

Biomedical subjects

Michael Ashburner

Publications and source records attributed to Michael Ashburner.

At least 19 recordsLinked to original sources

Getting a buzz out of the bee genome.

The honey bee Apis mellifera displays the most complex behavior of any insect. This, and its utility to humans, makes it a fascinating object of study for biologists. Such studies are now further enabled by the release of the honey-bee genome sequence.

Animals↗

EGASP: the human ENCODE Genome Annotation Assessment Project.

BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a 'reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40% to 50% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.

Alternative Splicing↗

National Center for Biomedical Ontology: advancing biomedicine through structured organization of scientific knowledge.

The National Center for Biomedical Ontology is a consortium that comprises leading informaticians, biologists, clinicians, and ontologists, funded by the National Institutes of Health (NIH) Roadmap, to develop innovative technology and methods that allow scientists to record, manage, and disseminate biomedical information and knowledge in machine-processable form. The goals of the Center are (1) to help unify the divergent and isolated efforts in ontology development by promoting high quality open-source, standards-based tools to create, manage, and use ontologies, (2) to create new software tools so that scientists can use ontologies to annotate and analyze biomedical data, (3) to provide a national resource for the ongoing evaluation, integration, and evolution of biomedical ontologies and associated tools and theories in the context of driving biomedical projects (DBPs), and (4) to disseminate the tools and resources of the Center and to identify, evaluate, and communicate best practices of ontology development to the biomedical community. Through the research activities within the Center, collaborations with the DBPs, and interactions with the biomedical community, our goal is to help scientists to work more effectively in the e-science paradigm, enhancing experiment design, experiment execution, data analysis, information synthesis, hypothesis generation and testing, and understand human disease.

Biomedical Research↗

Recurrent insertion and duplication generate networks of transposable element sequences in the Drosophila melanogaster genome.

BACKGROUND: The recent availability of genome sequences has provided unparalleled insights into the broad-scale patterns of transposable element (TE) sequences in eukaryotic genomes. Nevertheless, the difficulties that TEs pose for genome assembly and annotation have prevented detailed, quantitative inferences about the contribution of TEs to genomes sequences. RESULTS: Using a high-resolution annotation of TEs in Release 4 genome sequence, we revise estimates of TE abundance in Drosophila melanogaster. We show that TEs are non-randomly distributed within regions of high and low TE abundance, and that pericentromeric regions with high TE abundance are mosaics of distinct regions of extreme and normal TE density. Comparative analysis revealed that this punctate pattern evolves jointly by transposition and duplication, but not by inversion of TE-rich regions from unsequenced heterochromatin. Analysis of genome-wide patterns of TE nesting revealed a 'nesting network' that includes virtually all of the known TE families in the genome. Numerous directed cycles exist among TE families in the nesting network, implying concurrent or overlapping periods of transpositional activity. CONCLUSION: Rapid restructuring of the genomic landscape by transposition and duplication has recently added hundreds of kilobases of TE sequence to pericentromeric regions in D. melanogaster. These events create ragged transitions between unique and repetitive sequences in the zone between euchromatic and beta-heterochromatic regions. Complex relationships of TE nesting in beta-heterochromatic regions raise the possibility of a co-suppression network that may act as a global surveillance system against the majority of TE families in D. melanogaster.

Animals↗

Combined evidence annotation of transposable elements in genome sequences.

Transposable elements (TEs) are mobile, repetitive sequences that make up significant fractions of metazoan genomes. Despite their near ubiquity and importance in genome and chromosome biology, most efforts to annotate TEs in genome sequences rely on the results of a single computational program, RepeatMasker. In contrast, recent advances in gene annotation indicate that high-quality gene models can be produced from combining multiple independent sources of computational evidence. To elevate the quality of TE annotations to a level comparable to that of gene models, we have developed a combined evidence-model TE annotation pipeline, analogous to systems used for gene annotation, by integrating results from multiple homology-based and de novo TE identification methods. As proof of principle, we have annotated "TE models" in Drosophila melanogaster Release 4 genomic sequences using the combined computational evidence derived from RepeatMasker, BLASTER, TBLASTX, all-by-all BLASTN, RECON, TE-HMM and the previous Release 3.1 annotation. Our system is designed for use with the Apollo genome annotation tool, allowing automatic results to be curated manually to produce reliable annotations. The euchromatic TE fraction of D. melanogaster is now estimated at 5.3% (cf. 3.86% in Release 3.1), and we found a substantially higher number of TEs (n = 6,013) than previously identified (n = 1,572). Most of the new TEs derive from small fragments of a few hundred nucleotides long and highly abundant families not previously annotated (e.g., INE-1). We also estimated that 518 TE copies (8.6%) are inserted into at least one other TE, forming a nest of elements. The pipeline allows rapid and thorough annotation of even the most complex TE models, including highly deleted and/or nested elements such as those often found in heterochromatic sequences. Our pipeline can be easily adapted to other genome sequences, such as those of the D. melanogaster heterochromatin or other species in the genus Drosophila.

Journal Article↗

The Sequence Ontology: a tool for the unification of genome annotations.

The Sequence Ontology (SO) is a structured controlled vocabulary for the parts of a genomic annotation. SO provides a common set of terms and definitions that will facilitate the exchange, analysis and management of genomic data. Because SO treats part-whole relationships rigorously, data described with it can become substrates for automated reasoning, and instances of sequence features described by the SO can be subjected to a group of logical operations termed extensional mereology operators.

Alternative Splicing↗

An ontology for cell types.

We describe an ontology for cell types that covers the prokaryotic, fungal, animal and plant worlds. It includes over 680 cell types. These cell types are classified under several generic categories and are organized as a directed acyclic graph. The ontology is available in the formats adopted by the Open Biological Ontologies umbrella and is designed to be used in the context of model organism genome and other biological databases. The ontology is freely available at http://obo.sourceforge.net/ and can be viewed using standard ontology visualization tools such as OBO-Edit and COBrA.

Animals↗

Drosophila melanogaster: a case study of a model genomic sequence and its consequences.

The sequencing and annotation of the Drosophila melanogaster genome, first published in 2000 through collaboration between Celera Genomics and the Drosophila Genome Projects, has provided a number of important contributions to genome research. By demonstrating the utility of methods such as whole-genome shotgun sequencing and genome annotation by a community "jamboree," the Drosophila genome established the precedents for the current paradigm used by most genome projects. Subsequent releases of the initial genome sequence have been improved by the Berkeley Drosophila Genome Project and annotated by FlyBase, the Drosophila community database, providing one of the highest-quality genome sequences and annotations for any organism. We discuss the impact of the growing number of genome sequences now available in the genus on current Drosophila research, and some of the biological questions that these resources will enable to be solved in the future.

Animals↗

Molecular characterization of the singed wings locus of Drosophila melanogaster.

BACKGROUND: Hormones frequently guide animal development via the induction of cascades of gene activities, whose products further amplify an initial hormonal stimulus. In Drosophila the transformation of the larva into the pupa and the subsequent metamorphosis to the adult stage is triggered by changes in the titer of the steroid hormone 20-hydroxyecdysone. singed wings (swi) is the only gene known in Drosophila melanogaster for which mutations specifically interrupt the transmission of the regulatory signal from early to late ecdysone inducible genes. RESULTS: We have characterized singed wings locus, showing it to correspond to EG:171E4.2 (CG3095). swi encodes a predicted 68.5-kDa protein that contains N-terminal histidine-rich and threonine-rich domains, a cysteine-rich C-terminal region and two leucine-rich repeats. The SWI protein has a close homolog in D. melanogaster, defining a new family of SWI-like proteins, and is conserved in D. pseudoobscura. A lethal mutation, swit476, shows a severe disruption of the ecdysone pathway and is a C>Y substitution in one of the two conserved CysXCys motifs that are common to SWI and the Drosophila Toll-4 protein. CONCLUSIONS: It is not entirely clear from the present molecular analysis how the SWI protein may function in the ecdysone induced cascade. Currently all predictions agree in that SWI is very unlikely to be a nuclear protein. Thus it probably exercises its control of "late" ecdysone genes indirectly. Apparently the genetic regulation of ecdysone signaling is much more complex then was previously anticipated.

Amino Acid Sequence↗

A new hybrid rescue allele in Drosophila melanogaster.

Crosses of Drosophila melanogaster females to males of its sibling species Drosophila simulans, Drosophila mauritiana and Drosophila sechellia produce no sons and daughters that are viable only at low temperatures. We describe here a novel rescue allele Df(1)EP307-1-2 isolated on the basis of its suppression of high temperature hybrid female lethality. Df(1)EP307-1-2 also rescues hybrid males to the pharate adult stage, the same stage at which it is lethal to D. melanogaster pure species males. Molecular analysis indicates that Df(1)EP307-1-2 is associated with a deletion of about 61 kb in the 9D region of the X chromosome. The structure of Df(1)EP307-1-2 suggests that it was formed by a process similar to P-element induced male recombination.

Alleles↗

Drosophila crinkled, mutations of which disrupt morphogenesis and cause lethality, encodes fly myosin VIIA.

Myosin VIIs provide motor function for a wide range of eukaryotic processes. We demonstrate that mutations in crinkled (ck) disrupt the Drosophila myosin VIIA heavy chain. The ck/myoVIIA protein is present at a low level throughout fly development and at the same level in heads, thoraxes, and abdomens. Severe ck alleles, likely to be molecular nulls, die as embryos or larvae, but all allelic combinations tested thus far yield a small fraction of adult "escapers" that are weak and infertile. Scanning electron microscopy shows that escapers have defects in bristles and hairs, indicating that this motor protein plays a role in the structure of the actin cytoskeleton. We generate a homology model for the structure of the ck/myosin VIIA head that indicates myosin VIIAs, like myosin IIs, have a spectrin-like, SH3 subdomain fronting their N terminus. In addition, we establish that the two myosin VIIA FERM repeats share high sequence similarity with only the first two subdomains of the three-lobed structure that is typical of canonical FERM domains. Nevertheless, the approximately 100 and approximately 75 amino acids that follow the first two lobes of the first and second FERM domains are highly conserved among myosin VIIs, suggesting that they compose a conserved myosin tail homology 7 (MyTH7) domain that may be an integral part of the FERM domain or may function independently of it. Together, our data suggest a key role for ck/myoVIIA in the formation of cellular projections and other actin-based functions required for viability.

Amino Acid Sequence↗

The DrosDel collection: a set of P-element insertions for generating custom chromosomal aberrations in Drosophila melanogaster.

We describe a collection of P-element insertions that have considerable utility for generating custom chromosomal aberrations in Drosophila melanogaster. We have mobilized a pair of engineered P elements, p[RS3] and p[RS5], to collect 3243 lines unambiguously mapped to the Drosophila genome sequence. The collection contains, on average, an element every 35 kb. We demonstrate the utility of the collection for generating custom chromosomal deletions that have their end points mapped, with base-pair resolution, to the genome sequence. The collection was generated in an isogenic strain, thus affording a uniform background for screens where sensitivity to genetic background is high. The entire collection, along with a computational and genetic toolbox for designing and generating custom deletions, is publicly available. Using the collection it is theoretically possible to generate >12,000 deletions between 1 bp and 1 Mb in size by simple eye color selection. In addition, a further 37,000 deletions, selectable by molecular screening, may be generated. We are now using the collection to generate a second-generation deficiency kit that is precisely mapped to the genome sequence.

Animals↗

A novel system of fertility rescue in Drosophila hybrids reveals a link between hybrid lethality and female sterility.

Hybrid daughters of crosses between Drosophila melanogaster females and males from the D. simulans species clade are fully viable at low temperature but have agametic ovaries and are thus sterile. We report here that mutations in the D. melanogaster gene Hybrid male rescue (Hmr), along with unidentified polymorphic factors, rescue this agametic phenotype in both D. melanogaster/D. simulans and D. melanogaster/D. mauritiana F(1) female hybrids. These hybrids produced small numbers of progeny in backcrosses, their low fecundity being caused by incomplete rescue of oogenesis as well as by zygotic lethality. F(1) hybrid males from these crosses remained fully sterile. Hmr(+) is the first Drosophila gene shown to cause hybrid female sterility. These results also suggest that, while there is some common genetic basis to hybrid lethality and female sterility in D. melanogaster, hybrid females are more sensitive to fertility defects than to lethality.

Animals↗

Assessment of genome-wide protein function classification for Drosophila melanogaster.

The functional classification of genes on a genome-wide scale is now in its infancy, and we make a first attempt to assess existing methods and identify sources of error. To this end, we compared two independent efforts for associating proteins with functions, one implemented by FlyBase and the other by PANTHER at Celera Genomics. Both methods make inferences based on sequence similarity and the available experimental evidence. However, they differ considerably in methodology and process. Overall, assuming that the systematic error across the two methods is relatively small, we find the protein-to-function association error rate of both the FlyBase and PANTHER methods to be <2%. The primary source of error for both methods appears to be simple human error. Although homology-based inference can certainly cause errors in annotation, our analysis indicates that the frequency of such errors is relatively small compared with the number of correct inferences. Moreover, these homology errors can be minimized by careful tree-based inference, such as that implemented in PANTHER. Often, functional associations are made by one method and not the other, indicating that one of the greatest challenges lies in improving the completeness of available ontology associations.

Animals↗

Annotation of the Drosophila melanogaster euchromatic genome: a systematic review.

BACKGROUND: The recent completion of the Drosophila melanogaster genomic sequence to high quality and the availability of a greatly expanded set of Drosophila cDNA sequences, aligning to 78% of the predicted euchromatic genes, afforded FlyBase the opportunity to significantly improve genomic annotations. We made the annotation process more rigorous by inspecting each gene visually, utilizing a comprehensive set of curation rules, requiring traceable evidence for each gene model, and comparing each predicted peptide to SWISS-PROT and TrEMBL sequences. RESULTS: Although the number of predicted protein-coding genes in Drosophila remains essentially unchanged, the revised annotation significantly improves gene models, resulting in structural changes to 85% of the transcripts and 45% of the predicted proteins. We annotated transposable elements and non-protein-coding RNAs as new features, and extended the annotation of untranslated (UTR) sequences and alternative transcripts to include more than 70% and 20% of genes, respectively. Finally, cDNA sequence provided evidence for dicistronic transcripts, neighboring genes with overlapping UTRs on the same DNA sequence strand, alternatively spliced genes that encode distinct, non-overlapping peptides, and numerous nested genes. CONCLUSIONS: Identification of so many unusual gene models not only suggests that some mechanisms for gene regulation are more prevalent than previously believed, but also underscores the complex challenges of eukaryotic gene prediction. At present, experimental data and human curation remain essential to generate high-quality genome annotations.

Animals↗

The transposable elements of the Drosophila melanogaster euchromatin: a genomics perspective.

BACKGROUND: Transposable elements are found in the genomes of nearly all eukaryotes. The recent completion of the Release 3 euchromatic genomic sequence of Drosophila melanogaster by the Berkeley Drosophila Genome Project has provided precise sequence for the repetitive elements in the Drosophila euchromatin. We have used this genomic sequence to describe the euchromatic transposable elements in the sequenced strain of this species. RESULTS: We identified 85 known and eight novel families of transposable element varying in copy number from one to 146. A total of 1,572 full and partial transposable elements were identified, comprising 3.86% of the sequence. More than two-thirds of the transposable elements are partial. The density of transposable elements increases an average of 4.7 times in the centromere-proximal regions of each of the major chromosome arms. We found that transposable elements are preferentially found outside genes; only 436 of 1,572 transposable elements are contained within the 61.4 Mb of sequence that is annotated as being transcribed. A large proportion of transposable elements is found nested within other elements of the same or different classes. Lastly, an analysis of structural variation from different families reveals distinct patterns of deletion for elements belonging to different classes. CONCLUSIONS: This analysis represents an initial characterization of the transposable elements in the Release 3 euchromatic genomic sequence of D. melanogaster for which comparison to the transposable elements of other organisms can begin to be made. These data have been made available on the Berkeley Drosophila Genome Project website for future analyses.

Animals↗

Systematic determination of patterns of gene expression during Drosophila embryogenesis.

BACKGROUND: Cell-fate specification and tissue differentiation during development are largely achieved by the regulation of gene transcription. RESULTS: As a first step to creating a comprehensive atlas of gene-expression patterns during Drosophila embryogenesis, we examined 2,179 genes by in situ hybridization to fixed Drosophila embryos. Of the genes assayed, 63.7% displayed dynamic expression patterns that were documented with 25,690 digital photomicrographs of individual embryos. The photomicrographs were annotated using controlled vocabularies for anatomical structures that are organized into a developmental hierarchy. We also generated a detailed time course of gene expression during embryogenesis using microarrays to provide an independent corroboration of the in situ hybridization results. All image, annotation and microarray data are stored in publicly available database. We found that the RNA transcripts of about 1% of genes show clear subcellular localization. Nearly all the annotated expression patterns are distinct. We present an approach for organizing the data by hierarchical clustering of annotation terms that allows us to group tissues that express similar sets of genes as well as genes displaying similar expression patterns. CONCLUSIONS: Analyzing gene-expression patterns by in situ hybridization to whole-mount embryos provides an extremely rich dataset that can be used to identify genes involved in developmental processes that have been missed by traditional genetic analysis. Systematic analysis of rigorously annotated patterns of gene expression will complement and extend the types of analyses carried out using expression microarrays.

Animals↗

A hat trick--Plasmodium, Anopheles and Homo.

The genomes of the malaria parasite, its vector and its host are now sequenced. This has been a tremendous scientific achievement. But will it offer hope to the millions who die from malaria each year? Yes, but only if combined with political will and social change.

Animals↗