PubMed Health⌕ Search

Biomedical subjects

Patrick Wincker

Publications and source records attributed to Patrick Wincker.

29 records · Page 2Linked to original sources

Genome evolution in yeasts.

Identifying the mechanisms of eukaryotic genome evolution by comparative genomics is often complicated by the multiplicity of events that have taken place throughout the history of individual lineages, leaving only distorted and superimposed traces in the genome of each living organism. The hemiascomycete yeasts, with their compact genomes, similar lifestyle and distinct sexual and physiological properties, provide a unique opportunity to explore such mechanisms. We present here the complete, assembled genome sequences of four yeast species, selected to represent a broad evolutionary range within a single eukaryotic phylum, that after analysis proved to be molecularly as diverse as the entire phylum of chordates. A total of approximately 24,200 novel genes were identified, the translation products of which were classified together with Saccharomyces cerevisiae proteins into about 4,700 families, forming the basis for interspecific comparisons. Analysis of chromosome maps and genome redundancies reveal that the different yeast lineages have evolved through a marked interplay between several distinct molecular mechanisms, including tandem gene repeat formation, segmental duplication, a massive genome duplication and extensive gene loss.

Chromosomes, Fungal↗

Numerous novel annotations of the human genome sequence supported by a 5'-end-enriched cDNA collection.

A collection of 90,000 human cDNA clones generated to increase the fraction of "full-length" cDNAs available was analyzed by sequence alignment on the human genome assembly. Five hundred fifty-two gene models not found in LocusLink, with coding regions of at least 300 bp, were defined by using this collection. Exon composition proposed for novel genes showed an average of 4.7 exons per gene. In 20% of the cases, at least half of the exons predicted for new genes coincided with evolutionary conserved regions defined by sequence comparisons with the pufferfish Tetraodon nigroviridis. Among this subset, CpG islands were observed at the 5' end of 75%. In-frame stop codons upstream of the initiator ATG were present in 49% of the new genes, and 16% contained a coding region comprising at least 50% of the cDNA sequence. This cDNA resource also provided candidate small protein-coding genes, usually not included in genome annotations. In addition, analysis of a sample from this cDNA collection indicates that approximately 380 gene models described in LocusLink could be extended at their 5' end by at least one new exon. Finally, this cDNA resource provided an experimental support for annotations based exclusively on predictions, thus representing a resource substantially improving the human genome annotation.

5' Untranslated Regions↗

Whole genome sequence comparisons and "full-length" cDNA sequences: a combined approach to evaluate and improve Arabidopsis genome annotation.

To evaluate the existing annotation of the Arabidopsis genome further, we generated a collection of evolutionary conserved regions (ecores) between Arabidopsis and rice. The ecore analysis provides evidence that the gene catalog of Arabidopsis is not yet complete, and that a number of these annotations require re-examination. To improve the Arabidopsis genome annotation further, we used a novel "full-length" enriched cDNA collection prepared from several tissues. An additional 1931 genes were covered by new "full-length" cDNA sequences, raising the number of annotated genes with a corresponding "full-length" cDNA sequence to about 14,000. Detailed comparisons between these "full-length" cDNA sequences and annotated genes show that this resource is very helpful in determining the correct structure of genes, in particular, those not yet supported by "full-length" cDNAs. In addition, a total of 326 genomic regions not included previously in the Arabidopsis genome annotation were detected by this cDNA resource, providing clues for new gene discovery. Because, as expected, the two data sets only partially overlap, their combination produces very useful information for improving the Arabidopsis genome annotation.

Arabidopsis↗

A panoramic view of gene expression in the human kidney.

To gain a molecular understanding of kidney functions, we established a high-resolution map of gene expression patterns in the human kidney. The glomerulus and seven different nephron segments were isolated by microdissection from fresh tissue specimens, and their transcriptome was characterized by using the serial analysis of gene expression (SAGE) method. More than 400,000 mRNA SAGE tags were sequenced, making it possible to detect in each structure transcripts present at 18 copies per cell with a 95% confidence level. Expression of genes responsible for nephron transport and permeability properties was evidenced through transcripts for 119 solute carriers, 84 channels, 43 ion-transport ATPases, and 12 claudins. Searching for differences between the transcriptomes, we found 998 transcripts greatly varying in abundance from one nephron portion to another. Clustering analysis of these transcripts evidenced different extents of similarity between the nephron portions. Approximately 75% of the differentially distributed transcripts corresponded to cDNAs of known or unknown function that are accurately mapped in the human genome. This systematic large-scale analysis of individual structures of a complex human tissue reveals sets of genes underlying the function of well-defined nephron portions. It also provides quantitative expression data for a variety of genes mutated in hereditary diseases and helps in sorting candidate genes for renal diseases that affect specific portions of the human nephron.

Cluster Analysis↗

Genome sequence of the cyanobacterium Prochlorococcus marinus SS120, a nearly minimal oxyphototrophic genome.

Prochlorococcus marinus, the dominant photosynthetic organism in the ocean, is found in two main ecological forms: high-light-adapted genotypes in the upper part of the water column and low-light-adapted genotypes at the bottom of the illuminated layer. P. marinus SS120, the complete genome sequence reported here, is an extremely low-light-adapted form. The genome of P. marinus SS120 is composed of a single circular chromosome of 1,751,080 bp with an average G+C content of 36.4%. It contains 1,884 predicted protein-coding genes with an average size of 825 bp, a single rRNA operon, and 40 tRNA genes. Together with the 1.66-Mbp genome of P. marinus MED4, the genome of P. marinus SS120 is one of the two smallest genomes of a photosynthetic organism known to date. It lacks many genes that are involved in photosynthesis, DNA repair, solute uptake, intermediary metabolism, motility, phototaxis, and other functions that are conserved among other cyanobacteria. Systems of signal transduction and environmental stress response show a particularly drastic reduction in the number of components, even taking into account the small size of the SS120 genome. In contrast, housekeeping genes, which encode enzymes of amino acid, nucleotide, cofactor, and cell wall biosynthesis, are all present. Because of its remarkable compactness, the genome of P. marinus SS120 might approximate the minimal gene complement of a photosynthetic organism.

Adaptation, Physiological↗

Identification of the Anopheles gambiae ATP-binding cassette transporter superfamily genes.

The Anopheles gambiae genome sequence has been analyzed to find ATP-binding cassette protein genes based on deduced protein similarity to known family members. A nonredundant collection of 44 putative genes was identified including five genes not detected by the original Anopheles genome project machine annotation. These genes encode at least one member of all the human and Drosophila melanogaster ATP-binding protein subgroups. Like D. melanogaster, A. gambiae has subgroup ABCH genes encoding proteins different from the ABC proteins found in other complex organisms. The largest Anopheles subgroup is the ABCC genes which includes one member that can potentially encode ten different isoforms of the protein by differential splicing. As with Drosophila, the second largest Anopheles group is the ABCG subgroup with 12 genes compared to 15 genes in D. melanogaster, but only 5 genes in the human genome. In contrast, fewer ABCA and ABCB genes were identified in the mosquito genome than in the human or Drosophila genomes. Gene duplication is very evident in the Anopheles ABC genes with two groups of four genes, one group with three genes and three groups with two head to tail duplicated genes. These characteristics argue that the A. gambiae is actively using gene duplication as a mechanism to drive genetic variation in this important gene group.

ATP-Binding Cassette Transporters↗

The complete mitochondrial genome sequence of the pathogenic yeast Candida (Torulopsis) glabrata.

We report here the complete sequence of the mitochondrial (mt) genome of the pathogenic yeast Candida glabrata. This 20 kb mt genome is the smallest among sequenced hemiascomycetous yeasts. Despite its compaction, the mt genome contains the genes encoding the apocytochrome b (COB), three subunits of ATP synthetase (ATP6, 8 and 9), three subunits of cytochrome oxidase (COX1, 2 and 3), the ribosomal protein VAR1, 23 tRNAs, small and large ribosomal RNAs and the RNA subunit of RNase P. Three group I introns each with an intronic open reading frame are present in the COX1 gene. This sequence is available under accession number AJ511533.

Adenosine Triphosphatases↗

The DNA sequence and analysis of human chromosome 14.

Chromosome 14 is one of five acrocentric chromosomes in the human genome. These chromosomes are characterized by a heterochromatic short arm that contains essentially ribosomal RNA genes, and a euchromatic long arm in which most, if not all, of the protein-coding genes are located. The finished sequence of human chromosome 14 comprises 87,410,661 base pairs, representing 100% of its euchromatic portion, in a single continuous segment covering the entire long arm with no gaps. Two loci of crucial importance for the immune system, as well as more than 60 disease genes, have been localized so far on chromosome 14. We identified 1,050 genes and gene fragments, and 393 pseudogenes. On the basis of comparisons with other vertebrate genomes, we estimate that more than 96% of the chromosome 14 genes have been annotated. From an analysis of the CpG island occurrences, we estimate that 70% of these annotated genes are complete at their 5' end.

5' Untranslated Regions↗

Assessing the Drosophila melanogaster and Anopheles gambiae genome annotations using genome-wide sequence comparisons.

We performed genome-wide sequence comparisons at the protein coding level between the genome sequences of Drosophila melanogaster and Anopheles gambiae. Such comparisons detect evolutionarily conserved regions (ecores) that can be used for a qualitative and quantitative evaluation of the available annotations of both genomes. They also provide novel candidate features for annotation. The percentage of ecores mapping outside annotations in the A. gambiae genome is about fourfold higher than in D. melanogaster. The A. gambiae genome assembly also contains a high proportion of duplicated ecores, possibly resulting from artefactual sequence duplications in the genome assembly. The occurrence of 4063 ecores in the D. melanogaster genome outside annotations suggests that some genes are not yet or only partially annotated. The present work illustrates the power of comparative genomics approaches towards an exhaustive and accurate establishment of gene models and gene catalogues in insect genomes.

Animals↗

The genome sequence of the malaria mosquito Anopheles gambiae.

Anopheles gambiae is the principal vector of malaria, a disease that afflicts more than 500 million people and causes more than 1 million deaths each year. Tenfold shotgun sequence coverage was obtained from the PEST strain of A. gambiae and assembled into scaffolds that span 278 million base pairs. A total of 91% of the genome was organized in 303 scaffolds; the largest scaffold was 23.1 million base pairs. There was substantial genetic variation within this strain, and the apparent existence of two haplotypes of approximately equal frequency ("dual haplotypes") in a substantial fraction of the genome likely reflects the outbred nature of the PEST strain. The sequence produced a conservative inference of more than 400,000 single-nucleotide polymorphisms that showed a markedly bimodal density distribution. Analysis of the genome sequence revealed strong evidence for about 14,000 protein-encoding transcripts. Prominent expansions in specific families of proteins likely involved in cell adhesion and immunity were noted. An expressed sequence tag analysis of genes regulated by blood feeding provided insights into the physiological adaptations of a hematophagous insect.

Animals↗

Comparative genome and proteome analysis of Anopheles gambiae and Drosophila melanogaster.

Comparison of the genomes and proteomes of the two diptera Anopheles gambiae and Drosophila melanogaster, which diverged about 250 million years ago, reveals considerable similarities. However, numerous differences are also observed; some of these must reflect the selection and subsequent adaptation associated with different ecologies and life strategies. Almost half of the genes in both genomes are interpreted as orthologs and show an average sequence identity of about 56%, which is slightly lower than that observed between the orthologs of the pufferfish and human (diverged about 450 million years ago). This indicates that these two insects diverged considerably faster than vertebrates. Aligned sequences reveal that orthologous genes have retained only half of their intron/exon structure, indicating that intron gains or losses have occurred at a rate of about one per gene per 125 million years. Chromosomal arms exhibit significant remnants of homology between the two species, although only 34% of the genes colocalize in small "microsyntenic" clusters, and major interarm transfers as well as intra-arm shuffling of gene order are detected.

Animals↗