PubMed Health⌕ Search

Biomedical subjects

A Düsterhöft

Publications and source records attributed to A Düsterhöft.

At least 19 recordsLinked to original sources

The genome sequence of Schizosaccharomyces pombe.

We have sequenced and annotated the genome of fission yeast (Schizosaccharomyces pombe), which contains the smallest number of protein-coding genes yet recorded for a eukaryote: 4,824. The centromeres are between 35 and 110 kilobases (kb) and contain related repeats including a highly conserved 1.8-kb element. Regions upstream of genes are longer than in budding yeast (Saccharomyces cerevisiae), possibly reflecting more-extended control regions. Some 43% of the genes contain introns, of which there are 4,730. Fifty genes have significant similarity with human disease genes; half of these are cancer related. We identify highly conserved genes important for eukaryotic cell organization including those required for the cytoskeleton, compartmentation, cell-cycle control, proteolysis, protein phosphorylation and RNA splicing. These genes may have originated with the appearance of eukaryotic life. Few similarly conserved genes that are important for multicellular organization were identified, suggesting that the transition from prokaryotes to eukaryotes required more new genes than did the transition from unicellular to multicellular organization.

Base Sequence↗

Complete genome sequence and comparative analysis of the metabolically versatile Pseudomonas putida KT2440.

Pseudomonas putida is a metabolically versatile saprophytic soil bacterium that has been certified as a biosafety host for the cloning of foreign genes. The bacterium also has considerable potential for biotechnological applications. Sequence analysis of the 6.18 Mb genome of strain KT2440 reveals diverse transport and metabolic systems. Although there is a high level of genome conservation with the pathogenic Pseudomonad Pseudomonas aeruginosa (85% of the predicted coding regions are shared), key virulence factors including exotoxin A and type III secretion systems are absent. Analysis of the genome gives insight into the non-pathogenic nature of P. putida and points to potential new applications in agriculture, biocatalysis, bioremediation and bioplastic production.

Bacterial Proteins↗

Conservation of microstructure between a sequenced region of the genome of rice and multiple segments of the genome of Arabidopsis thaliana.

The nucleotide sequence was determined for a 340-kb segment of rice chromosome 2, revealing 56 putative protein-coding genes. This represents a density of one gene per 6.1 kb, which is higher than was reported for a previously sequenced segment of the rice genome. Sixteen of the putative genes were supported by matches to ESTs. The predicted products of 29 of the putative genes showed similarity to known proteins, and a further 17 genes showed similarity only to predicted or hypothetical proteins identified in genome sequence data. The region contains a few transposable elements: one retrotransposon, and one transposon. The segment of the rice genome studied had previously been identified as representing a part of rice chromosome 2 that may be homologous to a segment of Arabidopsis chromosome 4. We confirmed the conservation of gene content and order between the two genome segments. In addition, we identified a further four segments of the Arabidopsis genome that contain conserved gene content and order. In total, 22 of the 56 genes identified in the rice genome segment were represented in this set of Arabidopsis genome segments, with at least five genes present, in conserved order, in each segment. These data are consistent with the hypothesis that the Arabidopsis genome has undergone multiple duplication events. Our results demonstrate that conservation of the genome microstructure can be identified even between monocot and dicot species. However, the frequent occurrence of duplication, and subsequent microstructure divergence, within plant genomes may necessitate the integration of subsets of genes present in multiple redundant segments to deduce evolutionary relationships and identify orthologous genes.

Arabidopsis↗

Toward a catalog of human genes and proteins: sequencing and analysis of 500 novel complete protein coding human cDNAs.

With the complete human genomic sequence being unraveled, the focus will shift to gene identification and to the functional analysis of gene products. The generation of a set of cDNAs, both sequences and physical clones, which contains the complete and noninterrupted protein coding regions of all human genes will provide the indispensable tools for the systematic and comprehensive analysis of protein function to eventually understand the molecular basis of man. Here we report the sequencing and analysis of 500 novel human cDNAs containing the complete protein coding frame. Assignment to functional categories was possible for 52% (259) of the encoded proteins, the remaining fraction having no similarities with known proteins. By aligning the cDNA sequences with the sequences of the finished chromosomes 21 and 22 we identified a number of genes that either had been completely missed in the analysis of the genomic sequences or had been wrongly predicted. Three of these genes appear to be present in several copies. We conclude that full-length cDNA sequencing continues to be crucial also for the accurate identification of genes. The set of 500 novel cDNAs, and another 1000 full-coding cDNAs of known transcripts we have identified, adds up to cDNA representations covering 2%--5 % of all human genes. We thus substantially contribute to the generation of a gene catalog, consisting of both full-coding cDNA sequences and clones, which should be made freely available and will become an invaluable tool for detailed functional studies.

3' Untranslated Regions↗

Sequence and analysis of chromosome 5 of the plant Arabidopsis thaliana.

The genome of the model plant Arabidopsis thaliana has been sequenced by an international collaboration, The Arabidopsis Genome Initiative. Here we report the complete sequence of chromosome 5. This chromosome is 26 megabases long; it is the second largest Arabidopsis chromosome and represents 21% of the sequenced regions of the genome. The sequence of chromosomes 2 and 4 have been reported previously and that of chromosomes 1 and 3, together with an analysis of the complete genome sequence, are reported in this issue. Analysis of the sequence of chromosome 5 yields further insights into centromere structure and the sequence determinants of heterochromatin condensation. The 5,874 genes encoded on chromosome 5 reveal several new functions in plants, and the patterns of gene organization provide insights into the mechanisms and extent of genome evolution in plants.

Animals↗

Progress in Arabidopsis genome sequencing and functional genomics.

Arabidopsis thaliana has a relatively small genome of approximately 130 Mb containing about 10% repetitive DNA. Genome sequencing studies reveal a gene-rich genome, predicted to contain approximately 25000 genes spaced on average every 4.5 kb. Between 10 to 20% of the predicted genes occur as clusters of related genes, indicating that local sequence duplication and subsequent divergence generates a significant proportion of gene families. In addition to gene families, repetitive sequences comprise individual and small clusters of two to three retroelements and other classes of smaller repeats. The clustering of highly repetitive elements is a striking feature of the A. thaliana genome emerging from sequence and other analyses.

Agriculture↗

Analysis of deletion phenotypes and GFP fusions of 21 novel Saccharomyces cerevisiae open reading frames.

As part of EUROFAN (European Functional Analysis Network), we investigated 21 novel yeast open reading frames (ORFs) by growth and sporulation tests of deletion mutants. Two genes (YNL026w and YNL075w) are essential for mitotic growth and three deletion strains (ynl080c, ynl081c and ynl225c) grew with reduced rates. Two genes (YNL223w and YNL225c) were identified to be required for sporulation. In addition we also performed green fluorescent protein (GFP) tagging for localization studies. GFP labelling indicated the spindle pole body (Ynl225c-GFP) and the nucleus (Ynl075w-GFP) as the sites of action of two proteins. Ynl080c-GFP and Ynl081c-GFP fluorescence was visible in dot-shaped and elongated structures, whereas the Ynl022c-GFP signal was always found as one spot per cell, usually in the vicinity of nuclear DNA. The remaining C-terminal GFP fusions did not produce a clearly identifiable fluorescence signal. For 10 ORFs we constructed 5'-GFP fusions that were expressed from the regulatable GAL1 promoter. In all cases we observed GFP fluorescence upon induction but the localization of the fusion proteins remained difficult to determine. GFP-Ynl020c and GFP-Ynl034w strains grew only poorly on galactose, indicating a toxic effect of the overexpressed fusion proteins. In summary, we obtained a discernible GFP localization pattern in five of 20 strains investigated (25%). A deletion phenotype was observed in seven of 21 (33%) and an overexpression phenotype in two of 10 (20%) cases.

Gene Deletion↗

Automated sample-preparation technologies in genome sequencing projects.

A robotic workstation system (BioRobot 96OO, QIAGEN) and a 96-well UV spectrophotometer (Spectramax 250, Molecular Devices) were integrated in to the process of high-throughput automated sequencing of double-stranded plasmid DNA templates. An automated 96-well miniprep kit protocol (QIAprep Turbo, QIAGEN) provided high-quality plasmid DNA from shotgun clones. The DNA prepared by this procedure was used to generate more than two mega bases of final sequence data for two genomic projects (Arabidopsis thaliana and Schizosaccharomyces pombe), three thousand expressed sequence tags (ESTs) plus half a mega base of human full-length cDNA clones, and approximately 53,000 single reads for a whole genome shotgun project (Pseudomonas putida).

Arabidopsis↗

Sequence and analysis of chromosome 4 of the plant Arabidopsis thaliana.

The higher plant Arabidopsis thaliana (Arabidopsis) is an important model for identifying plant genes and determining their function. To assist biological investigations and to define chromosome structure, a coordinated effort to sequence the Arabidopsis genome was initiated in late 1996. Here we report one of the first milestones of this project, the sequence of chromosome 4. Analysis of 17.38 megabases of unique sequence, representing about 17% of the genome, reveals 3,744 protein coding genes, 81 transfer RNAs and numerous repeat elements. Heterochromatic regions surrounding the putative centromere, which has not yet been completely sequenced, are characterized by an increased frequency of a variety of repeats, new repeats, reduced recombination, lowered gene density and lowered gene expression. Roughly 60% of the predicted protein-coding genes have been functionally characterized on the basis of their homology to known genes. Many genes encode predicted proteins that are homologous to human and Caenorhabditis elegans proteins.

Animals↗

Introns and intein coding sequence in the ribonucleotide reductase genes of Bacillus subtilis temperate bacteriophage SPbeta.

The two putative ribonucleotide reductase subunits of the Bacillus subtilis bacteriophage SPbeta are encoded by the bnrdE and bnrdF genes that are highly similar to corresponding host paralogs, located on the opposite replication arm. In contrast to their bacterial counterparts, bnrdE and bnrdF each are interrupted by a group I intron, efficiently removed in vivo by mRNA processing. The bnrdF intron contains an ORF encoding a polypeptide similar to homing endonucleases responsible for intron mobility, whereas the bnrdE intron has no obvious trace of coding sequence. The downstream bnrdE exon harbors an intervening sequence not excised at the level of the primary transcript, which encodes an in-frame polypeptide displaying all the features of an intein. Presently, this is the only intein identified in bacteriophages. In addition, bnrdE provides an example of a group I intron and an intein coding sequence within the same gene.

Amino Acid Sequence↗

Analysis of 1.9 Mb of contiguous sequence from chromosome 4 of Arabidopsis thaliana.

The plant Arabidopsis thaliana (Arabidopsis) has become an important model species for the study of many aspects of plant biology. The relatively small size of the nuclear genome and the availability of extensive physical maps of the five chromosomes provide a feasible basis for initiating sequencing of the five chromosomes. The YAC (yeast artificial chromosome)-based physical map of chromosome 4 was used to construct a sequence-ready map of cosmid and BAC (bacterial artificial chromosome) clones covering a 1.9-megabase (Mb) contiguous region, and the sequence of this region is reported here. Analysis of the sequence revealed an average gene density of one gene every 4.8 kilobases (kb), and 54% of the predicted genes had significant similarity to known genes. Other interesting features were found, such as the sequence of a disease-resistance gene locus, the distribution of retroelements, the frequent occurrence of clustered gene families, and the sequence of several classes of genes not previously encountered in plants.

Arabidopsis↗

High-throughput robotic system for sequencing of microbial genomes.

A high-throughput robotic workstation system was used for double-stranded plasmid DNA template preparation and sequencing reaction setup to streamline the sequencing process in genome projects. All 96-well miniprep kits that were tested provided high quality plasmid DNA suitable for fluorescent DNA sequencing. After quantitation in a 96-well UV spectrophotometer, the plasmid DNA was used as template to automatically set up sequencing reactions. The setup was controlled by spread sheets that were imported into the robotic system. We utilized this integrated system to prepare all necessary shotgun templates for our contributions to a number of large-scale genome projects as well as a full-length cDNA sequencing project.

Arabidopsis↗

The complete genome sequence of the gram-positive bacterium Bacillus subtilis.

Bacillus subtilis is the best-characterized member of the Gram-positive bacteria. Its genome of 4,214,810 base pairs comprises 4,100 protein-coding genes. Of these protein-coding genes, 53% are represented once, while a quarter of the genome corresponds to several gene families that have been greatly expanded by gene duplication, the largest family containing 77 putative ATP-binding transport proteins. In addition, a large proportion of the genetic capacity is devoted to the utilization of a variety of carbon sources, including many plant-derived molecules. The identification of five signal peptidase genes, as well as several genes for components of the secretion apparatus, is important given the capacity of Bacillus strains to secrete large amounts of industrially important enzymes. Many of the genes are involved in the synthesis of secondary metabolites, including antibiotics, that are more typically associated with Streptomyces species. The genome contains at least ten prophages or remnants of prophages, indicating that bacteriophage infection has played an important evolutionary role in horizontal gene transfer, in particular in the propagation of bacterial pathogenesis.

Bacillus subtilis↗

The nucleotide sequence of Saccharomyces cerevisiae chromosome XII.

The yeast Saccharomyces cerevisiae is the pre-eminent organism for the study of basic functions of eukaryotic cells. All of the genes of this simple eukaryotic cell have recently been revealed by an international collaborative effort to determine the complete DNA sequence of its nuclear genome. Here we describe some of the features of chromosome XII.

Base Sequence↗

The nucleotide sequence of Saccharomyces cerevisiae chromosome XIV and its evolutionary implications.

In 1992 we started assembling an ordered library of cosmid clones from chromosome XIV of the yeast Saccharomyces cerevisiae. At that time, only 49 genes were known to be located on this chromosome and we estimated that 80% to 90% of its genes were yet to be discovered. In 1993, a team of 20 European laboratories began the systematic sequence analysis of chromosome XIV. The completed and intensively checked final sequence of 784,328 base pairs was released in April, 1996. Substantial parts had been published before or had previously been made available on request. The sequence contained 419 known or presumptive protein-coding genes, including two pseudogenes and three retrotransposons, 14 tRNA genes, and three small nuclear RNA genes. For 116 (30%) protein-coding sequences, one or more structural homologues were identified elsewhere in the yeast genome. Half of them belong to duplicated groups of 6-14 loosely linked genes, in most cases with conserved gene order and orientation (relaxed interchromosomal synteny). We have considered the possible evolutionary origins of this unexpected feature of yeast genome organization.

Base Sequence↗

Completion of the Saccharomyces cerevisiae genome sequence allows identification of KTR5, KTR6 and KTR7 and definition of the nine-membered KRE2/MNT1 mannosyltransferase gene family in this organism.

The KRE2/MNT1 mannosyltransferase gene family of Saccharomyces cerevisiae currently consists of the KRE2, YUR1, KTR1, KTR2, KTR3 and KTR4 genes. All six encode putative type II membrane proteins with a short cytoplasmic N-terminus, a membrane-spanning region and a highly conserved catalytic lumenal domain. Here we report the identification of the three remaining members of this family in the yeast genome. KTR5 corresponds to an open reading frame (ORF) of the left arm of chromosome XIV, and KTR6 and KTR7 to ORFs on the left arms of chromosomes XVI and IX respectively. The KTR5, KTR6 and KTR7 gene products are highly similar to the Kre2p/Mnt1p family members. Initial functional characterization revealed that some mutant yeast strains containing null copies of these genes displayed cell wall phenotypes. None was K1 killer toxin resistant but ktr6 and ktr7 null mutants were found to be hypersensitive and resistant, respectively, to the drug Calcofluor White.

Amino Acid Sequence↗

Allelism of PSO4 and PRP19 links pre-mRNA processing with recombination and error-prone DNA repair in Saccharomyces cerevisiae.

The radiation-sensitive mutant pso4-1 of Saccharomyces cerevisiae shows a pleiotropic phenotype, including sensitivity to DNA cross-linking agents, nearly blocked sporulation and reduced mutability. We have cloned the putative yeast DNA repair gene PSO4 from a genomic library by complementation of the blocked UV-induced mutagenesis and of sporulation in diploids homozygous for pso4-1. Sequence analysis revealed that gene PSO4 consists of 1512 bp located upstream of UBI4 on chromosome XII and encodes a putative protein of 56.7 kDa. PSO4 is allelic to PRP19, a gene encoding a spliceosome-associated protein, but shares no significant homology with other yeast genes. Gene disruption with a destroyed reading frame of our PSO4 clone resulted in death of haploid cells, confirming the finding that PSO4/PRP19 is an essential gene. Thus, PSO4 is the third essential DNA repair gene found in the yeast S.cerevisiae.

Alleles↗

Stepwise assembly of the lipid-linked oligosaccharide in the endoplasmic reticulum of Saccharomyces cerevisiae: identification of the ALG9 gene encoding a putative mannosyl transferase.

The core oligosaccharide Glc3Man9GlcNAc2 is assembled at the membrane of the endoplasmic reticulum on the lipid carrier dolichyl pyrophosphate and transferred to selected asparagine residues of nascent polypeptide chains. This transfer is catalyzed by the oligosaccharyl transferase complex. Based on the synthetic phenotype of the oligosaccharyl transferase mutation wbp1 in combination with a deficiency in the assembly pathway of the oligosaccharide in Saccharomyces cerevisiae, we have identified the novel ALG9 gene. We conclude that this locus encodes a putative mannosyl transferase because deletion of the gene led to accumulation of lipid-linked Man6GlcNAc2 in vivo and to hypoglycosylation of secreted proteins. Using an approach combining genetic and biochemical techniques, we show that the assembly of the lipid-linked core oligosaccharide in the lumen of the endoplasmic reticulum occurs in a stepwise fashion.

Amino Acid Sequence↗