PubMed Health⌕ Search

Biomedical subjects

Bruce Birren

Publications and source records attributed to Bruce Birren.

12 recordsLinked to original sources

Patterns of intron gain and loss in fungi.

Little is known about the patterns of intron gain and loss or the relative contributions of these two processes to gene evolution. To investigate the dynamics of intron evolution, we analyzed orthologous genes from four filamentous fungal genomes and determined the pattern of intron conservation. We developed a probabilistic model to estimate the most likely rates of intron gain and loss giving rise to these observed conservation patterns. Our data reveal the surprising importance of intron gain. Between about 150 and 250 gains and between 150 and 350 losses were inferred in each lineage. We discuss one gene in particular (encoding 1-phosphoribosyl-5-pyrophosphate synthetase) that displays an unusually high rate of intron gain in multiple lineages. It has been recognized that introns are biased towards the 5' ends of genes in intron-poor genomes but are evenly distributed in intron-rich genomes. Current models attribute this bias to 3' intron loss through a poly-adenosine-primed reverse transcription mechanism. Contrary to standard models, we find no increased frequency of intron loss toward the 3' ends of genes. Thus, recent intron dynamics do not support a model whereby 5' intron positional bias is generated solely by 3'-biased intron loss.

Adenosine↗

Genome duplication in the teleost fish Tetraodon nigroviridis reveals the early vertebrate proto-karyotype.

Tetraodon nigroviridis is a freshwater puffer fish with the smallest known vertebrate genome. Here, we report a draft genome sequence with long-range linkage and substantial anchoring to the 21 Tetraodon chromosomes. Genome analysis provides a greatly improved fish gene catalogue, including identifying key genes previously thought to be absent in fish. Comparison with other vertebrates and a urochordate indicates that fish proteins have diverged markedly faster than their mammalian homologues. Comparison with the human genome suggests approximately 900 previously unannotated human genes. Analysis of the Tetraodon and human genomes shows that whole-genome duplication occurred in the teleost fish lineage, subsequent to its divergence from mammals. The analysis also makes it possible to infer the basic structure of the ancestral bony vertebrate genome, which was composed of 12 chromosomes, and to reconstruct much of the evolutionary history of ancient and recent chromosome rearrangements leading to the modern human karyotype.

Animals↗

Genomic representations using concatenates of Type IIB restriction endonuclease digestion fragments.

We have developed a method for genomic representation using Type IIB restriction endonucleases. Representation by concatenation of restriction digests, or RECORD, is an approach to sample the fragments generated by cleavage with these enzymes. Here, we show that the RECORD libraries may be used for digital karyotyping and for pathogen identification by computational subtraction.

Bacteria↗

Methods in comparative genomics: genome correspondence, gene identification and regulatory motif discovery.

In Kellis et al. (2003), we reported the genome sequences of S. paradoxus, S. mikatae, and S. bayanus and compared these three yeast species to their close relative, S. cerevisiae. Genomewide comparative analysis allowed the identification of functionally important sequences, both coding and noncoding. In this companion paper we describe the mathematical and algorithmic results underpinning the analysis of these genomes. (1) We present methods for the automatic determination of genome correspondence. The algorithms enabled the automatic identification of orthologs for more than 90% of genes and intergenic regions across the four species despite the large number of duplicated genes in the yeast genome. The remaining ambiguities in the gene correspondence revealed recent gene family expansions in regions of rapid genomic change. (2) We present methods for the identification of protein-coding genes based on their patterns of nucleotide conservation across related species. We observed the pressure to conserve the reading frame of functional proteins and developed a test for gene identification with high sensitivity and specificity. We used this test to revisit the genome of S. cerevisiae, reducing the overall gene count by 500 genes (10% of previously annotated genes) and refining the gene structure of hundreds of genes. (3) We present novel methods for the systematic de novo identification of regulatory motifs. The methods do not rely on previous knowledge of gene function and in that way differ from the current literature on computational motif discovery. Based on genomewide conservation patterns of known motifs, we developed three conservation criteria that we used to discover novel motifs. We used an enumeration approach to select strongly conserved motif cores, which we extended and collapsed into a small number of candidate regulatory motifs. These include most previously known regulatory motifs as well as several noteworthy novel motifs. The majority of discovered motifs are enriched in functionally related genes, allowing us to infer a candidate function for novel motifs. Our results demonstrate the power of comparative genomics to further our understanding of any species. Our methods are validated by the extensive experimental knowledge in yeast and will be invaluable in the study of complex genomes like that of the human.

Algorithms↗

The complete genome and proteome of Mycoplasma mobile.

Although often considered "minimal" organisms, mycoplasmas show a wide range of diversity with respect to host environment, phenotypic traits, and pathogenicity. Here we report the complete genomic sequence and proteogenomic map for the piscine mycoplasma Mycoplasma mobile, noted for its robust gliding motility. For the first time, proteomic data are used in the primary annotation of a new genome, providing validation of expression for many of the predicted proteins. Several novel features were discovered including a long repeating unit of DNA of approximately 2435 bp present in five complete copies that are shown to code for nearly identical yet uniquely expressed proteins. M. mobile has among the lowest DNA GC contents (24.9%) and most reduced set of tRNAs of any organism yet reported (28). Numerous instances of tandem duplication as well as lateral gene transfer are evident in the genome. The multiple available complete genome sequences for other motile and immotile mycoplasmas enabled us to use comparative genomic and phylogenetic methods to suggest several candidate genes that might be involved in motility. The results of these analyses leave open the possibility that gliding motility might have arisen independently more than once in the mycoplasma lineage.

Amino Acid Sequence↗

Sequencing and comparison of yeast species to identify genes and regulatory elements.

Identifying the functional elements encoded in a genome is one of the principal challenges in modern biology. Comparative genomics should offer a powerful, general approach. Here, we present a comparative analysis of the yeast Saccharomyces cerevisiae based on high-quality draft sequences of three related species (S. paradoxus, S. mikatae and S. bayanus). We first aligned the genomes and characterized their evolution, defining the regions and mechanisms of change. We then developed methods for direct identification of genes and regulatory motifs. The gene analysis yielded a major revision to the yeast gene catalogue, affecting approximately 15% of all genes and reducing the total count by about 500 genes. The motif analysis automatically identified 72 genome-wide elements, including most known regulatory motifs and numerous new motifs. We inferred a putative function for most of these motifs, and provided insights into their combinatorial interactions. The results have implications for genome analysis of diverse organisms, including the human.

Base Sequence↗

The genome sequence of the filamentous fungus Neurospora crassa.

Neurospora crassa is a central organism in the history of twentieth-century genetics, biochemistry and molecular biology. Here, we report a high-quality draft sequence of the N. crassa genome. The approximately 40-megabase genome encodes about 10,000 protein-coding genes--more than twice as many as in the fission yeast Schizosaccharomyces pombe and only about 25% fewer than in the fruitfly Drosophila melanogaster. Analysis of the gene set yields insights into unexpected aspects of Neurospora biology including the identification of genes potentially associated with red light photobiology, genes implicated in secondary metabolism, and important differences in Ca2+ signalling as compared with plants and animals. Neurospora possesses the widest array of genome defence mechanisms known for any eukaryotic organism, including a process unique to fungi called repeat-induced point mutation (RIP). Genome analysis suggests that RIP has had a profound impact on genome evolution, greatly slowing the creation of new genes through genomic duplication and resulting in a genome with an unusually low proportion of closely related genes.

Calcium Signaling↗

Pathogen discovery from human tissue by sequence-based computational subtraction.

We have recently reported a new pathogen discovery approach, "computational subtraction". With this approach, non-human transcripts are detected by sequencing cDNA libraries from infected tissue and eliminating those transcripts that match the human genome. We show now that this method is experimentally feasible. We generated a cDNA library from a tissue sample of post-transplant lymphoproliferative disorder (PTLD). 27,840 independent cDNA sequences were filtered by computational subtraction against the known human sequence to identify 32 nonmatching transcripts. Of these, 22 (0.1%) were found to be amplifiable from both infected and noninfected samples and were inferred to be human DNA not yet contained in the available human genome sequence. The remaining 10 sequences could be amplified only from Epstein-Barr virus (EBV)-infected tissues. All 10 corresponded to the known EBV sequence. This proof-of-principle experiment demonstrates that computational subtraction can detect pathogenic microbes in primary human-diseased tissue.

DNA, Complementary↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗

The genome of M. acetivorans reveals extensive metabolic and physiological diversity.

Methanogenesis, the biological production of methane, plays a pivotal role in the global carbon cycle and contributes significantly to global warming. The majority of methane in nature is derived from acetate. Here we report the complete genome sequence of an acetate-utilizing methanogen, Methanosarcina acetivorans C2A. Methanosarcineae are the most metabolically diverse methanogens, thrive in a broad range of environments, and are unique among the Archaea in forming complex multicellular structures. This diversity is reflected in the genome of M. acetivorans. At 5,751,492 base pairs it is by far the largest known archaeal genome. The 4524 open reading frames code for a strikingly wide and unanticipated variety of metabolic and cellular capabilities. The presence of novel methyltransferases indicates the likelihood of undiscovered natural energy sources for methanogenesis, whereas the presence of single-subunit carbon monoxide dehydrogenases raises the possibility of nonmethanogenic growth. Although motility has not been observed in any Methanosarcineae, a flagellin gene cluster and two complete chemotaxis gene clusters were identified. The availability of genetic methods, coupled with its physiological and metabolic diversity, makes M. acetivorans a powerful model organism for the study of archaeal biology. [Sequence, data, annotations and analyses are available at http://www-genome.wi.mit.edu/.]

Archaeal Proteins↗

Lymphopenia in the BB rat model of type 1 diabetes is due to a mutation in a novel immune-associated nucleotide (Ian)-related gene.

The BB (BioBreeding) rat is one of the best models of spontaneous autoimmune diabetes and is used to study non-MHC loci contributing to Type 1 diabetes. Type 1 diabetes in the diabetes-prone BB (BBDP) rat is polygenic, dependent upon mutations at several loci. Iddm1, on chromosome 4, is responsible for a lymphopenia (lyp) phenotype and is essential to diabetes. In this study, we report the positional cloning of the Iddm1/lyp locus. We show that lymphopenia is due to a frameshift deletion in a novel member (Ian5) of the Immune-Associated Nucleotide (IAN)-related gene family, resulting in truncation of a significant portion of the protein. This mutation was absent in 37 other inbred rat strains that are nonlymphopenic and nondiabetic. The IAN gene family, lying within a tight cluster on rat chromosome 4, mouse chromosome 6, and human chromosome 7, is poorly characterized. Some members of the family have been shown to be expressed in mature T cells and switched on during thymic T-cell development, suggesting that Ian5 may be a key factor in T-cell development. The lymphopenia mutation may thus be useful not only to elucidate Type 1 diabetes, but also in the function of the Ian gene family as a whole.

Amino Acid Sequence↗

Structure and evolution of the Smith-Magenis syndrome repeat gene clusters, SMS-REPs.

An approximately 4-Mb genomic segment on chromosome 17p11.2, commonly deleted in patients with the Smith-Magenis syndrome (SMS) and duplicated in patients with dup(17)(p11.2p11.2) syndrome, is flanked by large, complex low-copy repeats (LCRs), termed proximal and distal SMS-REP. A third copy, the middle SMS-REP, is located between them. SMS-REPs are believed to mediate nonallelic homologous recombination, resulting in both SMS deletions and reciprocal duplications. To delineate the genomic structure and evolutionary origin of SMS-REPs, we constructed a bacterial artificial chromosome/P1 artificial chromosome contig spanning the entire SMS region, including the SMS-REPs, determined its genomic sequence, and used fluorescence in situ hybridization to study the evolution of SMS-REP in several primate species. Our analysis shows that both the proximal SMS-REP (approximately 256 kb) and the distal copy (approximately 176 kb) are located in the same orientation and derived from a progenitor copy, whereas the middle SMS-REP (approximately 241 kb) is inverted and appears to have been derived from the proximal copy. The SMS-REP LCRs are highly homologous (>98%) and contain at least 14 genes/pseudogenes each. SMS-REPs are not present in mice and were duplicated after the divergence of New World monkeys from pre-monkeys approximately 40-65 million years ago. Our findings potentially explain why the vast majority of SMS deletions and dup(17)(p11.2p11.2) occur at proximal and distal SMS-REPs and further support previous observations that higher-order genomic architecture involving LCRs arose recently during primate speciation and may predispose the human genome to both meiotic and mitotic rearrangements.

Abnormalities, Multiple↗