PubMed Health⌕ Search

Biomedical subjects

Kenneth H Wolfe

Publications and source records attributed to Kenneth H Wolfe.

At least 19 recordsLinked to original sources

Complete DNA sequences of the mitochondrial genomes of the pathogenic yeasts Candida orthopsilosis and Candida metapsilosis: insight into the evolution of linear DNA genomes from mitochondrial telomere mutants.

We determined complete mitochondrial DNA sequences of the two yeast species, Candida orthopsilosis and Candida metapsilosis, and compared them with the linear mitochondrial genome of their close relative, C.parapsilosis. Mitochondria of all the three species harbor compact genomes encoding the same set of genes arranged in the identical order. Differences in the length of these genomes result mainly from the presence/absence of introns. Multiple alterations were identified also in the sequences of the ribosomal and transfer RNAs, and proteins. However, the most striking feature of C.orthopsilosis and C.metapsilosis is the existence of strains differing in the molecular form of the mitochondrial genome (circular-mapping versus linear). Their analysis opens a unique window for understanding the role of mitochondrial telomeres in the stability and evolution of molecular architecture of the genome. Our results indicate that the circular-mapping mitochondrial genome derived from the linear form by intramolecular end-to-end fusions. Moreover, we suggest that the linear mitochondrial genome evolved from a circular-mapping form present in a common ancestor of the three species and, at the same time, the emergence of mitochondrial telomeres enabled the formation of linear monomeric DNA forms. In addition, comparison of isogenic C.metapsilosis strains differing in the form of the organellar genome suggests a possibility that, under some circumstances, the linearity and/or the presence of telomeres provide a competitive advantage over a circular-mapping mitochondrial genome.

Base Sequence↗

Functional partitioning of yeast co-expression networks after genome duplication.

Several species of yeast, including the baker's yeast Saccharomyces cerevisiae, underwent a genome duplication roughly 100 million years ago. We analyze genetic networks whose members were involved in this duplication. Many networks show detectable redundancy and strong asymmetry in their interactions. For networks of co-expressed genes, we find evidence for network partitioning whereby the paralogs appear to have formed two relatively independent subnetworks from the ancestral network. We simulate the degeneration of networks after duplication and find that a model wherein the rate of interaction loss depends on the "neighborliness" of the interacting genes produces networks with parameters similar to those seen in the real partitioned networks. We propose that the rationalization of network structure through the loss of pair-wise gene interactions after genome duplication provides a mechanism for the creation of semi-independent daughter networks through the division of ancestral functions between these daughter networks.

Base Sequence↗

Comparative genomics and genome evolution in yeasts.

Yeasts provide a powerful model system for comparative genomics research. The availability of multiple complete genome sequences from different fungal groups--currently 18 hemiascomycetes, 8 euascomycetes and 4 basidiomycetes--enables us to gain a broad perspective on genome evolution. The sequenced genomes span a continuum of divergence levels ranging from multiple individuals within a species to species pairs with low levels of protein sequence identity and no conservation of gene order. One of the most interesting emerging areas is the growing number of events such as gene losses, gene displacements and gene relocations that can be attributed to the action of natural selection.

Centromere↗

Multiple rounds of speciation associated with reciprocal gene loss in polyploid yeasts.

A whole-genome duplication occurred in a shared ancestor of the yeast species Saccharomyces cerevisiae, Saccharomyces castellii and Candida glabrata. Here we trace the subsequent losses of duplicated genes, and show that the pattern of loss differs among the three species at 20% of all loci. For example, several transcription factor genes, including STE12, TEC1, TUP1 and MCM1, are single-copy in S. cerevisiae but are retained in duplicate in S. castellii and C. glabrata. At many loci, different species have lost different members of a duplicated gene pair, so that 4-7% of single-copy genes compared between any two species are not orthologues. This pattern of gene loss provides strong evidence for speciation through a version of the Bateson-Dobzhansky-Muller mechanism, in which the loss of alternative copies of duplicated genes leads to reproductive isolation. We show that the lineages leading to the three species diverged shortly after the whole-genome duplication, during a period of precipitous gene loss. The set of loci at which single-copy paralogues are retained is biased towards genes involved in ribosome biogenesis and genes that evolve slowly, consistent with the hypothesis that reciprocal gene loss is more likely to occur between duplicated genes that are functionally indistinguishable. We propose a simple, unified model in which a single mechanism--passive gene loss-enabled whole--genome duplication and led to the rapid emergence of new yeast species.

Alleles↗

Visualizing syntenic relationships among the hemiascomycetes with the Yeast Gene Order Browser.

The Yeast Gene Order Browser (YGOB) is an online tool designed to facilitate the comparative genomic visualization and appraisal of synteny within and between the genomes of seven hemiascomycete yeast species. Three of these genomes are polyploid, and hence contain intra-genomic syntenic regions, the correct assembly of which is a particular success of YGOB. Designed to accurately assemble, display and score gene order relationships, YGOB is both an interactive tool for browsing genomic data, and a software engine now being used for evolutionary analyses on a whole-genome scale. Underlying the online interface is the YGOB database, which consists of homology assignments across the species, extensively curated based on sequence similarity and novelly, an appraisal of genomic context (synteny) in multiple genomes. Currently the YGOB database incorporates genome data from Saccharomyces cerevisiae, Candida glabrata, Saccharomyces castellii, Ashbya gossypii, Kluyveromyces lactis, Kluyveromyces waltii and Saccharomyces kluyveri, but the system is scaleable to accommodate additional genomes. This paper discusses the usage and utility of version 1.0 of YGOB, which is publicly available at http://wolfe.gen.tcd.ie/ygob.

Chromosomes, Fungal↗

Gene duplication, exon gain and neofunctionalization of OEP16-related genes in land plants.

OEP16, a channel protein of the outer membrane of chloroplasts, has been implicated in amino acid transport and in the substrate-dependent import of protochlorophyllide oxidoreductase A. Two major clades of OEP16-related sequences were identified in land plants (OEP16-L and OEP16-S), which arose by a gene duplication event predating the divergence of seed plants and bryophytes. Remarkably, in angiosperms, OEP16-S genes evolved by gaining an additional exon that extends an interhelical loop domain in the pore-forming region of the protein. We analysed the sequence, structure and expression of the corresponding Arabidopsis genes (atOEP16-S and atOEP16-L) and demonstrated that following duplication, both genes diverged in terms of expression patterns and coding sequence. AtOEP16-S, which contains multiple G-box ABA-responsive elements (ABREs) in the promoter region, is regulated by ABI3 and ABI5 and is strongly expressed during the maturation phase in seeds and pollen grains, both desiccation-tolerant tissues. In contrast, atOEP-L, which lacks promoter ABREs, is expressed predominantly in leaves, is induced strongly by low-temperature stress and shows weak induction in response to osmotic stress, salicylic acid and exogenous ABA. Our results indicate that gene duplication, exon gain and regulatory sequence evolution each played a role in the divergence of OEP16 homologues in plants.

Amino Acid Sequence↗

Rate asymmetry after genome duplication causes substantial long-branch attraction artifacts in the phylogeny of Saccharomyces species.

Whole-genome duplication (WGD) produces sets of gene pairs that are all of the same age. We therefore expect that phylogenetic trees that relate these pairs to their orthologs in other species should show a single consistent topology. However, a previous study of gene pairs formed by WGD in the yeast Saccharomyces cerevisiae found conflicting topologies among neighbor-joining (NJ) trees drawn from different loci and suggested that this conflict was the result of "asynchronous functional divergence" of duplicated genes (Langkjaer, R. B., P. F. Cliften, M. Johnston, and J. Piskur. 2003. Yeast genome duplication was followed by asynchronous differentiation of duplicated genes. Nature 421:848-852). Here, we test whether the conflicting topologies might instead be due to asymmetrical rates of evolution leading to long-branch attraction (LBA) artifacts in phylogenetic trees. We constructed trees for 433 pairs of WGD paralogs in S. cerevisiae with their single orthologs in Saccharomyces kluyveri and Candida albicans. We find a strong correlation between the asymmetry of evolutionary rates of a pair of S. cerevisiae paralogs and the topology of the tree inferred for that pair. Saccharomyces cerevisiae gene pairs with approximately equal rates of evolution tend to give phylogenies in which the WGD postdates the speciation between S. cerevisiae and S. kluyveri (B-trees), whereas trees drawn from gene pairs with asymmetrical rates tend to show WGD pre-dating this speciation (A-trees). Gene order data from throughout the genome indicate that the "A-trees" are artifacts, even though more than 50% of gene pairs are inferred to have this topology when the NJ method as implemented in ClustalW (i.e., with Poisson correction of distances) is used to construct the trees. This LBA artifact can be ameliorated, but not eliminated, by using gamma-corrected distances or by using maximum likelihood trees with robustness estimated by the Shimodaira-Hasegawa test. Tests for adaptive evolution indicated that positive selection might be the cause of rate asymmetry in a substantial fraction (19%) of the paralog pairs.

Evolution, Molecular↗

The Yeast Gene Order Browser: combining curated homology and syntenic context reveals gene fate in polyploid species.

We developed the Yeast Gene Order Browser (YGOB; http://wolfe.gen.tcd.ie/ygob) to facilitate visual comparisons and computational analysis of synteny relationships in yeasts. The data presented in YGOB, currently covering seven species, are based on sets of homologous genes that have been intensively manually curated based on both sequence similarity and genomic context (synteny). We reconciled different laboratories' lists of paralogous Saccharomyces cerevisiae gene pairs formed by genome duplication (ohnologs), and present near-exhaustive lists of the ohnolog pairs retained in S. cerevisiae (551, including 22 previously unidentified), Saccharomyces castellii (599), and Candida glabrata (404).

Databases, Genetic↗

Changes in alternative splicing of human and mouse genes are accompanied by faster evolution of constitutive exons.

Alternative splicing is known to be an important source of protein sequence variation, but its evolutionary impact has not been explored in detail. Studying alternative splicing requires extensive sampling of the transcriptome, but new data sets based on expressed sequence tags aligned to chromosomes make it possible to study alternative splicing on a genome-wide scale. Although genes showing alternative splicing by exon skipping are conserved as compared to the genome as a whole, we find that genes where structural differences between human and mouse result in genome-specific alternatively spliced exons in one species show almost 60% greater nonsynonymous divergence in constitutive exons than genes where exon skipping is conserved. This effect is also seen for genes showing species-specific patterns of alternative splicing where gene structure is conserved. Our observations are not attributable to an inherent difference in rate of evolution between these two sets of proteins or to differences with respect to predictors of evolutionary rate such as expression level, tissue specificity, or genetic redundancy. Where genome-specific alternatively spliced exons are seen in mammals, the vast majority of skipped exons appear to be recent additions to gene structures. Furthermore, among genes with genome-specific alternatively spliced exons, the degree of nonsynonymous divergence in constitutive sequence is a function of the frequency of incorporation of these alternative exons into transcripts. These results suggest that alterations in alternative splicing pattern can have knock-on effects in terms of accelerated sequence evolution in constant regions of the protein.

Alternative Splicing↗

Birth of a metabolic gene cluster in yeast by adaptive gene relocation.

Although most eukaryotic genomes lack operons, they contain some physical clusters of genes that are related in function despite being unrelated in sequence. How these clusters are formed during evolution is unknown. The DAL cluster is the largest metabolic gene cluster in yeast and consists of six adjacent genes encoding proteins that enable Saccharomyces cerevisiae to use allantoin as a nitrogen source. We show here that the DAL cluster was assembled, quite recently in evolutionary terms, through a set of genomic rearrangements that happened almost simultaneously. Six of the eight genes involved in allantoin degradation, which were previously scattered around the genome, became relocated to a single subtelomeric site in an ancestor of S. cerevisiae and Saccharomyces castellii. These genomic rearrangements coincided with a biochemical reorganization of the purine degradation pathway, which switched to importing allantoin instead of urate. This change eliminated urate oxidase, one of several oxygen-consuming enzymes that were lost by yeasts that can grow vigorously in anaerobic conditions. The DAL cluster is located in a domain of modified chromatin involving both H2A.Z histone exchange and Hst1-Sum1-mediated histone deacetylation, and it may be a coadapted gene complex formed by epistatic selection.

Allantoin↗

A genome sequence survey shows that the pathogenic yeast Candida parapsilosis has a defective MTLa1 allele at its mating type locus.

Candida parapsilosis is responsible for ca. 15% of Candida infections and is of particular concern in neonates and surgical intensive care patients. The related species Candida albicans has recently been shown to possess a functional mating pathway. To analyze the analogous pathway in C. parapsilosis, we carried out a genome sequence survey of the type strain. We identified ca. 3,900 genes, with an average amino acid identity of 59% with C. albicans. Of these, 23 are predicted to be predominantly involved in mating. We identified a genomic locus homologous to the MTLa mating type locus of C. albicans, but the C. parapsilosis type strain has at least two internal stop codons in the MTLa1 open reading frame, and two predicted introns are not spliced. These stop codons were present in MTLa1 of all eight C. parapsilosis isolates tested. Furthermore, we found that all isolates of C. parapsilosis tested appear to contain only the MTLa idiomorph at the presumptive mating locus, unlike C. albicans and C. dubliniensis. MTLalpha sequences are present but at a different chromosomal location. It is therefore likely that all (or at least the majority) of C. parapsilosis isolates have a mating pathway that is either defective or substantially different from that of C. albicans.

Alleles↗

Clusters of co-expressed genes in mammalian genomes are conserved by natural selection.

Genes that belong to the same functional pathways are often packaged into operons in prokaryotes. However, aside from examples in nematode genomes, this form of transcriptional regulation appears to be absent in eukaryotes. Nevertheless, a number of recent studies have shown that gene order in eukaryotic genomes is not completely random, and that genes with similar expression patterns tend to be clustered together. What remains unclear is whether co-expressed genes have been gathered together by natural selection to facilitate their regulation, or if the genes are co-expressed simply by virtue of their being close together in the genome. Here, we show that gene expression clusters tend to contain fewer chromosomal breakpoints between human and mouse than expected by chance, which indicates that they are being held together by natural selection. This conclusion applies to clusters defined on the basis of broad (housekeeping) expression, or on the basis of correlated transcription profiles across tissues. Contrary to previous reports, we find that genes with high expression are not clustered to a greater extent than expected by chance and are not conserved during evolution.

Animals↗

Allele-specific transcript isoforms in human.

Estimates of the number of human genes that produce more than one transcript isoform through alternative mRNA splicing depend on the assumption that the observation of multiple transcripts from a gene can be attributed entirely to alternative splicing. It is possible, however, that a substantial proportion of cases where multiple transcripts have been observed for a gene result from differences between alleles. Many examples of genes that are spliced differently from different alleles have been reported but no systematic estimate of the proportion of alternatively spliced genes that are affected by such polymorphisms has been carried out. We find that alternative transcript isoforms are non-randomly associated with closely linked nucleotide polymorphisms, based on an integrated analysis of the dbSNP, dbEST and ASAP databases. From the observed level of association between transcript isoforms and polymorphisms, we estimate that 21% of alternatively spliced genes are affected by polymorphisms that either completely determine which form of the transcript is observed or alter the relative abundances of some of the alternative isoforms. We provide a conservative lower bound of 6% on this estimate and point out that alternative splicing cannot be confirmed absolutely unless more than one transcript is observed from the same allele.

Alleles↗

Origins of recently gained introns in Caenorhabditis.

The genomes of the nematodes Caenorhabditis elegans and Caenorhabditis briggsae both contain approximately 100,000 introns, of which >6,000 are unique to one or the other species. To study the origins of new introns, we used a conservative method involving phylogenetic comparisons to animal orthologs and nematode paralogs to identify cases where an intron content difference between C. elegans and C. briggsae was caused by intron insertion rather than deletion. We identified 81 recently gained introns in C. elegans and 41 in C. briggsae. Novel introns have a stronger exon splice site consensus sequence than the general population of introns and show the same preference for phase 0 sites in codons over phases 1 and 2. More of the novel introns are inserted in genes that are expressed in the C. elegans germ line than expected by chance. Thirteen of the 122 gained introns are in genes whose protein products function in premRNA processing, including three gains in the gene for spliceosomal protein SF3B1 and two in the nonsense-mediated decay gene smg-2. Twenty-eight novel introns have significant DNA sequence identity to other introns, including three that are similar to other introns in the same gene. All of these similarities involve minisatellites or palindromes in the intron sequences. Our results suggest that at least some of the intron gains were caused by reverse splicing of a preexisting intron.

Amino Acid Sequence↗

PubCrawler: keeping up comfortably with PubMed and GenBank.

The free PubCrawler web service (http://www.pubcrawler.ie) has been operating for five years and so far has brought literature and sequence updates to over 22 000 users. It provides information on a personalized web page whenever new articles appear in PubMed or when new sequences are found in GenBank that are specific to customized queries. The server also acts as an automatic alerting system by sending out short notifications or emails with the latest updates as soon as they become available. A new output format and more flexibility for the email formatting help PubCrawler cope with increasing challenges arising from browser incompatibilities and mail filters, therefore making it suitable for a wide range of users.

Databases, Nucleic Acid↗

Widespread paleopolyploidy in model plant species inferred from age distributions of duplicate genes.

It is often anticipated that many of today's diploid plant species are in fact paleopolyploids. Given that an ancient large-scale duplication will result in an excess of relatively old duplicated genes with similar ages, we analyzed the timing of duplication of pairs of paralogous genes in 14 model plant species. Using EST contigs (unigenes), we identified pairs of paralogous genes in each species and used the level of synonymous nucleotide substitution to estimate the relative ages of gene duplication. For nine of the investigated species (wheat [Triticum aestivum], maize [Zea mays], tetraploid cotton [Gossypium hirsutum], diploid cotton [G. arboretum], tomato [Lycopersicon esculentum], potato [Solanum tuberosum], soybean [Glycine max], barrel medic [Medicago truncatula], and Arabidopsis thaliana), the age distributions of duplicated genes contain peaks corresponding to short evolutionary periods during which large numbers of duplicated genes were accumulated. Large-scale duplications (polyploidy or aneuploidy) are strongly suspected to be the cause of these temporal peaks of gene duplication. However, the unusual age profile of tandem gene duplications in Arabidopsis indicates that other scenarios, such as variation in the rate at which duplicated genes are deleted, must also be considered.

Computational Biology↗

Functional divergence of duplicated genes formed by polyploidy during Arabidopsis evolution.

To study the evolutionary effects of polyploidy on plant gene functions, we analyzed functional genomics data for a large number of duplicated gene pairs formed by ancient polyploidy events in Arabidopsis thaliana. Genes retained in duplicate are not distributed evenly among Gene Ontology or Munich Information Center for Protein Sequences functional categories, which indicates a nonrandom process of gene loss. Genes involved in signal transduction and transcription have been preferentially retained, and those involved in DNA repair have been preferentially lost. Although the two members of each gene pair must originally have had identical transcription profiles, less than half of the pairs formed by the most recent polyploidy event still retain significantly correlated profiles. We identified several cases where groups of duplicated gene pairs have diverged in concert, forming two parallel networks, each containing one member of each gene pair. In these cases, the expression of each gene is strongly correlated with the other nonhomologous genes in its network but poorly correlated with its paralog in the other network. We also find that the rate of protein sequence evolution has been significantly asymmetric in >20% of duplicate pairs. Together, these results suggest that functional diversification of the surviving duplicated genes is a major feature of the long-term evolution of polyploids.

Arabidopsis↗

Evolution of the MAT locus and its Ho endonuclease in yeast species.

The genetics of the mating-type (MAT) locus have been studied extensively in Saccharomyces cerevisiae, but relatively little is known about how this complex system evolved. We compared the organization of MAT and mating-type-like (MTL) loci in nine species spanning the hemiascomycete phylogenetic tree. We inferred that the system evolved in a two-step process in which silent HMR/HML cassettes appeared, followed by acquisition of the Ho endonuclease from a mobile genetic element. Ho-mediated switching between an active MAT locus and silent cassettes exists only in the Saccharomyces sensu stricto group and their closest relatives: Candida glabrata, Kluyveromyces delphensis, and Saccharomyces castellii. We identified C. glabrata MTL1 as the ortholog of the MAT locus of K. delphensis and show that switching between C. glabrata MTL1a and MTL1alpha genotypes occurs in vivo. The more distantly related species Kluyveromyces lactis has silent cassettes but switches mating type without the aid of Ho endonuclease. Very distantly related species such as Candida albicans and Yarrowia lipolytica do not have silent cassettes. In Pichia angusta, a homothallic species, we found MATalpha2, MATalpha1, and MATa1 genes adjacent to each other on the same chromosome. Although some continuity in the chromosomal location of the MAT locus can be traced throughout hemiascomycete evolution and even to Neurospora, the gene content of the locus has changed with the loss of an HMG domain gene (MATa2) from the MATa idiomorph shortly after HO was recruited.

Deoxyribonucleases, Type II Site-Specific↗