PubMed HealthSearch

Biomedical subjects

K Heumann

Publications and source records attributed to K Heumann.

9 recordsLinked to original sources

MIPS: a database for genomes and protein sequences.

The Munich Information Center for Protein Sequences (MIPS-GSF), Martinsried near Munich, Germany, develops and maintains genome oriented databases. It is commonplace that the amount of sequence data available increases rapidly, but not the capacity of qualified manual annotation at the sequence databases. Therefore, our strategy aims to cope with the data stream by the comprehensive application of analysis tools to sequences of complete genomes, the systematic classification of protein sequences and the active support of sequence analysis and functional genomics projects. This report describes the systematic and up-to-date analysis of genomes (PEDANT), a comprehensive database of the yeast genome (MYGD), a database reflecting the progress in sequencing the Arabidopsis thaliana genome (MATD), the database of assembled, annotated human EST clusters (MEST), and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). MIPS provides access through its WWW server (http://www.mips.biochem.mpg.de) to a spectrum of generic databases, including the above mentioned as well as a database of protein families (PROTFAM), the MITOP database, and the all-against-all FASTA database.

Amino Acid Sequence

Comprehensive, comprehensible, distributed and intelligent databases: current status.

MOTIVATION: It is only a matter of time until a user will see not many but one integrated database of information for molecular biology. Is this true? Is it a good thing? Why will it happen? Where are we now? What developments are fostering and what developments are impeding progress towards this end? SUPPLEMENTARY INFORMATION: A list of WWW resources devoted to database issues in molecular biology is available at http://www.mips.biochem.mpg.de CONTACT: frishman@mips.biochem.mpg.de

Computational Biology

Overview of the yeast genome.

The collaboration of more than 600 scientists from over 100 laboratories to sequence the Saccharomyces cerevisiae genome was the largest decentralised experiment in modern molecular biology and resulted in a unique data resource representing the first complete set of genes from a eukaryotic organism. 12 million bases were sequenced in a truly international effort involving European, US, Canadian and Japanese laboratories. While the yeast genome represents only a small fraction of the information in today's public sequence databases, the complete, ordered and non-redundant sequence provides an invaluable resource for the detailed analysis of cellular gene function and genome architecture. In terms of throughput, completeness and information content, yeast has always been the lead eukaryotic organism in genomics; it is still the largest genome to be completely sequenced.

Chromosome Mapping

The nucleotide sequence of Saccharomyces cerevisiae chromosome XII.

The yeast Saccharomyces cerevisiae is the pre-eminent organism for the study of basic functions of eukaryotic cells. All of the genes of this simple eukaryotic cell have recently been revealed by an international collaborative effort to determine the complete DNA sequence of its nuclear genome. Here we describe some of the features of chromosome XII.

Base Sequence

The nucleotide sequence of Saccharomyces cerevisiae chromosome XIV and its evolutionary implications.

In 1992 we started assembling an ordered library of cosmid clones from chromosome XIV of the yeast Saccharomyces cerevisiae. At that time, only 49 genes were known to be located on this chromosome and we estimated that 80% to 90% of its genes were yet to be discovered. In 1993, a team of 20 European laboratories began the systematic sequence analysis of chromosome XIV. The completed and intensively checked final sequence of 784,328 base pairs was released in April, 1996. Substantial parts had been published before or had previously been made available on request. The sequence contained 419 known or presumptive protein-coding genes, including two pseudogenes and three retrotransposons, 14 tRNA genes, and three small nuclear RNA genes. For 116 (30%) protein-coding sequences, one or more structural homologues were identified elsewhere in the yeast genome. Half of them belong to duplicated groups of 6-14 loosely linked genes, in most cases with conserved gene order and orientation (relaxed interchromosomal synteny). We have considered the possible evolutionary origins of this unexpected feature of yeast genome organization.

Base Sequence

MIPS: a database for protein sequences, homology data and yeast genome information.

The MIPS group (Martinsried Institute for Protein Sequences) at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, collects, processes and distributes protein sequence data within the framework of the tripartite association of the PIR-International Protein Sequence Database (,). MIPS contributes nearly 50% of the data input to the PIR-International Protein Sequence Database. The database is distributed on CD-ROM together with PATCHX, an exhaustive supplement of unique, unverified protein sequences from external sources compiled by MIPS. Through its WWW server (http://www.mips.biochem.mpg.de/ ) MIPS permits internet access to sequence databases, homology data and to yeast genome information. (i) Sequence similarity results from the FASTA program () are stored in the FASTA database for all proteins from PIR-International and PATCHX. The database is dynamically maintained and permits instant access to FASTA results. (ii) Starting with FASTA database queries, proteins have been classified into families and superfamilies (PROT-FAM). (iii) The HPT (hashed position tree) data structure () developed at MIPS is a new approach for rapid sequence and pattern searching. (iv) MIPS provides access to the sequence and annotation of the complete yeast genome (), the functional classification of yeast genes (FunCat) and its graphical display, the 'Genome Browser' (). A CD-ROM based on the JAVA programming language providing dynamic interactive access to the yeast genome and the related protein sequences has been compiled and is available on request.

Academies and Institutes

Complete nucleotide sequence of Saccharomyces cerevisiae chromosome X.

The complete nucleotide sequence of Saccharomyces cerevisiae chromosome X (745 442 bp) reveals a total of 379 open reading frames (ORFs), the coding region covering approximately 75% of the entire sequence. One hundred and eighteen ORFs (31%) correspond to genes previously identified in S. cerevisiae. All other ORFs represent novel putative yeast genes, whose function will have to be determined experimentally. However, 57 of the latter subset (another 15% of the total) encode proteins that show significant analogy to proteins of known function from yeast or other organisms. The remaining ORFs, exhibiting no significant similarity to any known sequence, amount to 54% of the total. General features of chromosome X are also reported, with emphasis on the nucleotide frequency distribution in the environment of the ATG and stop codons, the possible coding capacity of at least some of the small ORFs (<100 codons) and the significance of 46 non-canonical or unpaired nucleotides in the stems of some of the 24 tRNA genes recognized on this chromosome.

Amino Acid Sequence

A top-down approach to whole genome visualization.

The investigation of large DNA contigs like complete chromosomes or genomes requires novel methods of data visualization. The complex information contained in a genome, particularly the relation of its individual genetic elements, needs to be accessible in a comprehensive, intelligent and intelligible manner. The yeast genome is expected to contain more than 6,000 Open Reading Frames (ORFs). As yet, the function of many of these ORFs has not been characterized satisfactorily. Also, many ORFs are found to have redundant copies elsewhere in the genome that originated from common ancestors. Other genetic elements (e.g. Tss, delta-elements, t-RNAs) are present in multiple copies. To visualize these relationships, a top-down "genome browser" is introduced that enables inspection of genomic data at different levels of abstraction (e.g. chromosomes, coding/non-coding regions, high/low levels of similarity). This novel tool is a key component for the integrated services approach to biological sequence data management (Heumann et al. 1995) and is accessible through the world wide web (WWW). This work demonstrates how the genome browser visualizes the results of an all-against-all comparison of the elements in the yeast genome as a graph. Interactive navigational queries across yeast chromosomes along the lines of sequence similarity open versatile options for the detailed investigation of genome properties. For sequence comparison the hashed position tree HPT (Mewes & Heumann 1995) is applied. Sequence similarity relationships are represented using the genome similarity graph (GSG) (Heumann & Mewes 1996c).

Chromosomes, Fungal

A new concept of sequence data distribution on wide area networks.

Accepted concepts in distributed applications design have been applied in the development of a network-based system for the synchronization of remote sequence database access sites by an incremental update mechanism. Computer hardware requirements, network bandwidth, and stability considerations make centralized access to essential computerized resources undesirable. A network model has been developed to distribute access over a collection of remotely situated computer centers. The formally independent database-access nodes join to form a heterogeneous, long distance, co-operating network that can compensate for the deficiencies of unstable network links thereby ensuring uninterrupted access to the resource. In order to guarantee consistency among these nodes, several distributed transaction protocols have been investigated; based on these results, a prototype system has been implemented. A layered software architecture makes the distributed transaction protocol transparent to the individual database system and the underlying network. Individual components of this network communicate by means of Remote Procedure Calls (RPCs). A prototype software system operates to synchronize up to data copies of the PIR-International Protein Sequence Database (Barker et al., 1993) at a number of different sites using the public Internet as the transport vehicle.

Computer Communication Networks