Genome-wide analyses based on comparative genomics.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to W Saurin.
Explore the source record for details and available documents.
Ralstonia solanacearum is a devastating, soil-borne plant pathogen with a global distribution and an unusually wide host range. It is a model system for the dissection of molecular determinants governing pathogenicity. We present here the complete genome sequence and its analysis of strain GMI1000. The 5.8-megabase (Mb) genome is organized into two replicons: a 3.7-Mb chromosome and a 2.1-Mb megaplasmid. Both replicons have a mosaic structure providing evidence for the acquisition of genes through horizontal gene transfer. Regions containing genetically mobile elements associated with the percentage of G+C bias may have an important function in genome evolution. The genome encodes many proteins potentially associated with a role in pathogenicity. In particular, many putative attachment factors were identified. The complete repertoire of type III secreted effector proteins can be studied. Over 40 candidates were identified. Comparison with other genomes suggests that bacterial plant pathogens and animal pathogens harbour distinct arrays of specialized type III-dependent effectors.
Explore the source record for details and available documents.
Microsporidia are obligate intracellular parasites infesting many animal groups. Lacking mitochondria and peroxysomes, these unicellular eukaryotes were first considered a deeply branching protist lineage that diverged before the endosymbiotic event that led to mitochondria. The discovery of a gene for a mitochondrial-type chaperone combined with molecular phylogenetic data later implied that microsporidia are atypical fungi that lost mitochondria during evolution. Here we report the DNA sequences of the 11 chromosomes of the approximately 2.9-megabase (Mb) genome of Encephalitozoon cuniculi (1,997 potential protein-coding genes). Genome compaction is reflected by reduced intergenic spacers and by the shortness of most putative proteins relative to their eukaryote orthologues. The strong host dependence is illustrated by the lack of genes for some biosynthetic pathways and for the tricarboxylic acid cycle. Phylogenetic analysis lends substantial credit to the fungal affiliation of microsporidia. Because the E. cuniculi genome contains genes related to some mitochondrial functions (for example, Fe-S cluster assembly), we hypothesize that microsporidia have retained a mitochondrion-derived organelle.
We report the construction of a tiling path of around 650 clones covering more than 99% of human chromosome 14. Clone overlap information to assemble the map was derived by comparing fully sequenced clones with a database of clone end sequences (sequence tag connector strategy). We selected homogeneously distributed seed points using an auxiliary high-resolution radiation hybrid map comprising 1,895 distinct positions. The high long-range continuity and low redundancy of the tiling path indicates that the sequence tag connector approach compares favourably with alternative mapping strategies.
A DNA sequencing program was applied to the small (<3 Mb) genome of the microsporidian Encephalitozoon cuniculi, an amitochondriate eukaryotic parasite of mammals, and the sequence of the smallest chromosome was determined. The approximately 224-kb E. cuniculi chromosome I exhibits a dyad symmetry characterized by two identical 37-kb subtelomeric regions which are divergently oriented and extend just downstream of the inverted copies of an 8-kb duplicated cluster of six genes. Each subtelomeric region comprises a single 16S-23S rDNA transcription unit, flanked by various tandemly repeated sequences, and ends with approximately 1 kb of heterogeneous telomeric repeats. The central (or core) region of the chromosome harbors a highly compact arrangement of 132 potential protein-coding genes plus two tRNA genes (one gene per 1.14 kb). Most genes occur as single copies with no identified introns. Of these putative genes, only 53 could be assigned to known functions. A number of genes from the transcription and translation machineries as well as from other cellular processes display characteristic eukaryotic signatures or are clearly eukaryote-specific.
The identification of molecular evolutionary mechanisms in eukaryotes is approached by a comparative genomics study of a homogeneous group of species classified as Hemiascomycetes. This group includes Saccharomyces cerevisiae, the first eukaryotic genome entirely sequenced, back in 1996. A random sequencing analysis has been performed on 13 different species sharing a small genome size and a low frequency of introns. Detailed information is provided in the 20 following papers. Additional tables available on websites describe the ca. 20000 newly identified genes. This wealth of data, so far unique among eukaryotes, allowed us to examine the conservation of chromosome maps, to identify the 'yeast-specific' genes, and to review the distribution of gene families into functional classes. This project conducted by a network of seven French laboratories has been designated 'Génolevures'.
The generation of sequencing data for the hemiascomycetous yeast random sequence tag project was performed using the procedures established at GENOSCOPE. These procedures include a series of protocols for the sequencing reactions, using infra-red labelled primers, performed on both ends of the plasmid inserts in the same reaction tube, and their analysis on automated DNA sequencers. They also include a package of computer programs aimed at detecting potential assignation errors, selecting good quality sequences and estimating their useful length.
We have analyzed the evolution of chromosome maps of Hemiascomycetes by comparing gene order and orientation of the 13 yeast species partially sequenced in this program with the genome map of Saccharomyces cerevisiae. From the analysis of nearly 8000 situations in which two distinct genes having homologs in S. cerevisiae could be identified on the sequenced inserts of another yeast species, we have quantified the loss of synteny, the frequency of single gene deletion and the occurrence of gene inversion. Traces of ancestral duplications in the genome of S. cerevisiae could be identified from the comparison with the other species that do not entirely coincide with those identified from the comparison of S. cerevisiae with itself. From such duplications and from the correlation observed between gene inversion and loss of synteny, a model is proposed for the molecular evolution of Hemiascomycetes. This model, which can possibly be extended to other eukaryotes, is based on the reiteration of events of duplication of chromosome segments, creating transient merodiploids that are subsequently resolved by single gene deletion events.
Comparisons of the 6213 predicted Saccharomyces cerevisiae open reading frame (ORF) products with sequences from organisms of other biological phyla differentiate genes commonly conserved in evolution from 'maverick' genes which have no homologue in phyla other than the Ascomycetes. We show that a majority of the 'maverick' genes have homologues among other yeast species and thus define a set of 1892 genes that, from sequence comparisons, appear 'Ascomycetes-specific'. We estimate, retrospectively, that the S. cerevisiae genome contains 5651 actual protein-coding genes, 50 of which were identified for the first time in this work, and that the present public databases contain 612 predicted ORFs that are not real genes. Interestingly, the sequences of the 'Ascomycetes-specific' genes tend to diverge more rapidly in evolution than that of other genes. Half of the 'Ascomycetes-specific' genes are functionally characterized in S. cerevisiae, and a few functional categories are over-represented in them.
We have evaluated the degree of gene redundancy in the nuclear genomes of 13 hemiascomycetous yeast species. Saccharomyces cerevisiae singletons and gene families appear generally conserved in these species as singletons and families of similar size, respectively. Variations of the number of homologues with respect to that expected affect from 7 to less than 24% of each genome. Since S. cerevisiae homologues represent the majority of the genes identified in the genomes studied, the overall degree of gene redundancy seems conserved across all species. This is best explained by a dynamic equilibrium resulting from numerous events of gene duplication and deletion rather than by a massive duplication event occurring in some lineages and not in others.
We explored the biological diversity of hemiascomycetous yeasts using a set of 22000 newly identified genes in 13 species through BLASTX searches. Genes without clear homologue in Saccharomyces cerevisiae appeared to be conserved in several species, suggesting that they were recently lost by S. cerevisiae. They often identified well-known species-specific traits. Cases of gene acquisition through horizontal transfer appeared to occur very rarely if at all. All identified genes were ascribed to functional classes. Functional classes were differently represented among species. Species classification by functional clustering roughly paralleled rDNA phylogeny. Unequal distribution of rapidly evolving, ascomycete-specific, genes among species and functions was shown to contribute strongly to this clustering. A few cases of gene family amplification were documented, but no general correlation could be observed between functional differentiation of yeast species and variations of gene family sizes. Yeast biological diversity seems thus to result from limited species-specific gene losses or duplications, and for a large part from rapid evolution of genes and regulatory factors dedicated to specific functions.
Arabidopsis thaliana is an important model system for plant biologists. In 1996 an international collaboration (the Arabidopsis Genome Initiative) was formed to sequence the whole genome of Arabidopsis and in 1999 the sequence of the first two chromosomes was reported. The sequence of the last three chromosomes and an analysis of the whole genome are reported in this issue. Here we present the sequence of chromosome 3, organized into four sequence segments (contigs). The two largest (13.5 and 9.2 Mb) correspond to the top (long) and the bottom (short) arms of chromosome 3, and the two small contigs are located in the genetically defined centromere. This chromosome encodes 5,220 of the roughly 25,500 predicted protein-coding genes in the genome. About 20% of the predicted proteins have significant homology to proteins in eukaryotic genomes for which the complete sequence is available, pointing to important conserved cellular functions among eukaryotes.
Despite a rapid increase in the amount of available archaeal sequence information, little is known about the duplication of genetic material in the third domain of life. We identified a single origin of bidirectional replication in Pyrococcus abyssi by means of in silico analyses of cumulative oligomer skew and the identification of an early replicating chromosomal segment. The replication origin in three Pyrococcus species was found to be highly conserved, and several eukaryotic-like DNA replication genes were clustered around it. As in Bacteria, the chromosomal region containing the replication terminus was a hot spot of genome shuffling. Thus, although bacterial and archaeal replication proteins differ profoundly, they are used to replicate chromosomes in a similar manner in both prokaryotic domains.
The number of genes in the human genome is unknown, with estimates ranging from 50,000 to 90,000 (refs 1, 2), and to more than 140,000 according to unpublished sources. We have developed 'Exofish', a procedure based on homology searches, to identify human genes quickly and reliably. This method relies on the sequence of another vertebrate, the pufferfish Tetraodon nigroviridis, to detect conserved sequences with a very low background. Similar to Fugu rubripes, a marine pufferfish proposed by Brenner et al. as a model for genomic studies, T. nigroviridis is a more practical alternative with a genome also eight times more compact than that of human. Many comparisons have been made between F. rubripes and human DNA that demonstrate the potential of comparative genomics using the pufferfish genome. Application of Exofish to the December version of the working draft sequence of the human genome and to Unigene showed that the human genome contains 28,000-34,000 genes, and that Unigene contains less than 40% of the protein-coding fraction of the human genome.
Tetraodon nigroviridis is a freshwater pufferfish 20-30 million years distant from Fugu rubripes. The genome of both tetraodontiforms is compact, mostly because intergenic and intronic sequences are reduced in size compared to other vertebrate genomes. The previously uncharacterized Tetraodon genome is described here together with a detailed analysis of its repeat content and organization. We report the sequencing of 46 megabases of bacterial artificial chromosome (BAC) end sequences, which represents a random DNA sample equivalent to 13% of the genome. The sequence and location of rRNA gene clusters, centromeric and subtelocentric satellite sequences have been determined. Minisatellites and microsatellites have been cataloged and notable differences were observed in comparison with microsatellites from Fugu. The genome contains homologies to all known families of transposable elements, including Ty3-gypsy, Ty1-copia, Line retrotransposons, DNA transposons, and retroviruses, although their overall abundance is <1%. This structural analysis is an important prerequisite to sequencing the Tetraodon genome.
ATP-binding cassette (ABC) systems, also called traffic ATPases, are found in eukaryotes and prokaryotes and almost all participate in the transport of a wide variety of molecules. ABC systems are characterized by a highly conserved ATPase module called here the ABC module, involved in coupling transport to ATP hydrolysis. We have used the sequence of one of the first representatives of bacterial ABC transporters, the MalK protein, to collect 250 closely related sequences from a nonredundant protein sequence database. The sequences collected by this objective method are all known or putative ABC transporters. After having eliminated short protein sequences and duplicates, the 197 remaining sequences were subjected to a phylogenetic analysis based on a mutational similarity matrix. An unrooted tree for these modules was found to display two major branches, one grouping all collected uptake systems and the other all collected export systems. This remarkable disposition strongly suggests that the divergence between these two functionally different types of ABC systems occurred once in the history of these systems and probably before the differentiation of prokaryotes and eukaryotes. We discuss the implications of this finding and we propose a model accounting for the generation and the diversification of ABC systems.
To address the mechanisms of host-virus adaptation and pathogenesis of lentiviral infections, we compared the evolution of the same isolate of simian immunodeficiency virus (SIVsmm9) in two different situations: nonpathogenic infection of its natural host, the sooty mangabey, and AIDS-inducing infection of a new host, the rhesus macaque. Samples were obtained at 6, 12, and 23 or 30 months postinfection from three animals of each species. Sequences were derived from the V1 and V2 domains of the surface glycoprotein. In the macaques, we observed specific variations absent from all mangabey samples, indicating that different host species select different virus variants. In the macaques, we also observed a different shape in the phylogenetic tree, a lower divergence of sibling sequences, and a lower synonymous/nonsynonymous change ratio than in the mangabeys. This suggests that the viral population is larger and submitted to weaker selection pressures when host-virus adaptation is achieved, such as in the mangabey.