PubMed Health⌕ Search

Biomedical subjects

Peer Bork

Publications and source records attributed to Peer Bork.

At least 37 records · Page 2Linked to original sources

Proteome survey reveals modularity of the yeast cell machinery.

Protein complexes are key molecular entities that integrate multiple gene products to perform cellular functions. Here we report the first genome-wide screen for complexes in an organism, budding yeast, using affinity purification and mass spectrometry. Through systematic tagging of open reading frames (ORFs), the majority of complexes were purified several times, suggesting screen saturation. The richness of the data set enabled a de novo characterization of the composition and organization of the cellular machinery. The ensemble of cellular proteins partitions into 491 complexes, of which 257 are novel, that differentially combine with additional attachment proteins or protein modules to enable a diversification of potential functions. Support for this modular organization of the proteome comes from integration with available data on expression, localization, function, evolutionary conservation, protein structure and binary interactions. This study provides the largest collection of physically determined eukaryotic cellular machines so far and a platform for biological data integration and modelling.

Genome, Fungal↗

LSAT: learning about alternative transcripts in MEDLINE.

MOTIVATION: Generation of alternative transcripts from the same gene is an important biological event due to their contribution in creating functional diversity in eukaryotes. In this work, we choose the task of extracting information around this complex topic using a two-step procedure involving machine learning and information extraction. RESULTS: In the first step, we trained a classifier that inductively learns to identify sentences about physiological transcript diversity from the MEDLINE abstracts. Using a large hand-built corpus, we compared the sentence classification performance of various text categorization methods. Support vector machines (SVMs) followed by the maximum entropy classifier outperformed other methods for the sentence classification task. The SVM with the radial basis function kernel and optimized parameters achieved Fbeta-measure of 91% during the 4-fold cross validation and of 74% when applied to all sentences in more than 12 million abstracts of MEDLINE. In the second step, we identified eight frequently present semantic categories in the sentences and performed a limited amount of semantic role labeling. The role labeling step also achieved very high Fbeta-measure for all eight categories. AVAILABILITY: The results of our two-step procedure are summarized in the LSAT database of alternative transcripts. LSAT is available at http://www.bork.embl.de/LSAT CONTACT: shah@embl.de SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Benchmarking↗

SMART 5: domains in the context of genomes and networks.

The Simple Modular Architecture Research Tool (SMART) is an online resource (http://smart.embl.de/) used for protein domain identification and the analysis of protein domain architectures. Many new features were implemented to make SMART more accessible to scientists from different fields. The new 'Genomic' mode in SMART makes it easy to analyze domain architectures in completely sequenced genomes. Domain annotation has been updated with a detailed taxonomic breakdown and a prediction of the catalytic activity for 50 SMART domains is now available, based on the presence of essential amino acids. Furthermore, intrinsically disordered protein regions can be identified and displayed. The network context is now displayed in the results page for more than 350 000 proteins, enabling easy analyses of domain interactions.

Catalysis↗

A temporal map of transcription factor activity: mef2 directly regulates target genes at all stages of muscle development.

Dissecting components of key transcriptional networks is essential for understanding complex developmental processes and phenotypes. Genetic studies have highlighted the role of members of the Mef2 family of transcription factors as essential regulators in myogenesis from flies to man. To understand how these transcription factors control diverse processes in muscle development, we have combined chromatin immunoprecipitation analysis with gene expression profiling to obtain a temporal map of Mef2 activity during Drosophila embryonic development. This global approach revealed three temporal patterns of Mef2 enhancer binding, providing a glimpse of dynamic enhancer use within the context of a developing embryo. Our results provide mechanistic insight into the regulation of Mef2's activity at the level of DNA binding and suggest cooperativity with the bHLH protein Twist. The number and diversity of new direct target genes indicates a much broader role for Mef2, at all stages of myogenesis, than previously anticipated.

Animals↗

Literature mining for the biologist: from information retrieval to biological discovery.

For the average biologist, hands-on literature mining currently means a keyword search in PubMed. However, methods for extracting biomedical facts from the scientific literature have improved considerably, and the associated tools will probably soon be used in many laboratories to automatically annotate and analyse the growing number of system-wide experimental data sets. Owing to the increasing body of text and the open-access policies of many journals, literature mining is also becoming useful for both hypothesis generation and biological discovery. However, the latter will require the integration of literature and high-throughput data, which should encourage close collaborations between biologists and computational linguists.

Computational Biology↗

Vertebrate-type intron-rich genes in the marine annelid Platynereis dumerilii.

Previous genome comparisons have suggested that one important trend in vertebrate evolution has been a sharp rise in intron abundance. By using genomic data and expressed sequence tags from the marine annelid Platynereis dumerilii, we provide direct evidence that about two-thirds of human introns predate the bilaterian radiation but were lost from insect and nematode genomes to a large extent. A comparison of coding exon sequences confirms the ancestral nature of Platynereis and human genes. Thus, the urbilaterian ancestor had complex, intron-rich genes that have been retained in Platynereis and human.

Animals↗

Spore number control and breeding in Saccharomyces cerevisiae: a key role for a self-organizing system.

Spindle pole bodies (SPBs) provide a structural basis for genome inheritance and spore formation during meiosis in yeast. Upon carbon source limitation during sporulation, the number of haploid spores formed per cell is reduced. We show that precise spore number control (SNC) fulfills two functions. SNC maximizes the production of spores (1-4) that are formed by a single cell. This is regulated by the concentration of three structural meiotic SPB components, which is dependent on available amounts of carbon source. Using experiments and computer simulation, we show that the molecular mechanism relies on a self-organizing system, which is able to generate particular patterns (different numbers of spores) in dependency on one single stimulus (gradually increasing amounts of SPB constituents). We also show that SNC enhances intratetrad mating, whereby maximal amounts of germinated spores are able to return to a diploid lifestyle without intermediary mitotic division. This is beneficial for the immediate fitness of the population of postmeiotic cells.

Carbon↗

Palindromic repetitive DNA elements with coding potential in Methanocaldococcus jannaschii.

We have identified 141 novel palindromic repetitive elements in the genome of euryarchaeon Methanocaldococcus jannaschii. The total length of these elements is 14.3kb, which corresponds to 0.9% of the total genomic sequence and 6.3% of all extragenic regions. The elements can be divided into three groups (MJRE1-3) based on the sequence similarity. The low sequence identity within each of the groups suggests rather old origin of these elements in M. jannaschii. Three MJRE2 elements were located within the protein coding regions without disrupting the coding potential of the host genes, indicating that insertion of repeats might be a widespread mechanism to enhance sequence diversity in coding regions.

Amino Acid Sequence↗

Medusa: a simple tool for interaction graph analysis.

SUMMARY: Medusa is a Java application for visualizing and manipulating graphs of interaction, such as data from the STRING database. It features an intuitive user interface developed with the help of biologists. Medusa is optimized for accessing protein interaction data from STRING, but can be used for any type of graph from any scientific field.

Algorithms↗

G2D: a tool for mining genes associated with disease.

BACKGROUND: Human inherited diseases can be associated by genetic linkage with one or more genomic regions. The availability of the complete sequence of the human genome allows examining those locations for an associated gene. We previously developed an algorithm to prioritize genes on a chromosomal region according to their possible relation to an inherited disease using a combination of data mining on biomedical databases and gene sequence analysis. RESULTS: We have implemented this method as a web application in our site G2D (Genes to Diseases). It allows users to inspect any region of the human genome to find candidate genes related to a genetic disease of their interest. In addition, the G2D server includes pre-computed analyses of candidate genes for 552 linked monogenic diseases without an associated gene, and the analysis of 18 asthma loci. CONCLUSION: G2D can be publicly accessed at http://www.ogic.ca/projects/g2d_2/.

Algorithms↗

Very-KIND is a novel nervous system specific guanine nucleotide exchange factor for Ras GTPases.

The kinase non-catalytic c-lobe domain (KIND) evolved from the catalytic protein kinase fold into a potential protein interaction module for signalling proteins. Spir family actin organizers and the non-receptor phosphatase type 13 (PTP type 13) encode a KIND domain in the very N-terminal parts of the proteins. Here we report the characterization and cloning of a third member of the KIND protein family, which we have named very-KIND (VKIND) because of its two KIND domains. Like the other members of the protein family, VKIND has a KIND domain at the N-terminus. A second KIND domain is located in the central part of the protein. The C-terminal half encodes a guanine nucleotide exchange factor motif for Ras-like GTPases (RasGEF) and a RasGEF N-terminal module (RasGEFN). There is only one VKIND gene in the mammalian genomes and up to now we have found the gene only in vertebrates. During mouse embryogenesis the VKIND gene was specifically expressed in the developing nervous system. In adult mice Northern hybridizations revealed high expression only in brain. Low expression could be detected in ovary. In situ hybridizations showed a specific expression of VKIND in neuronal cells of the granular and Purkinje cell layers of the cerebellum.

Amino Acid Sequence↗

Extraction of regulatory gene/protein networks from Medline.

MOTIVATION: We have previously developed a rule-based approach for extracting information on the regulation of gene expression in yeast. The biomedical literature, however, contains information on several other equally important regulatory mechanisms, in particular phosphorylation, which we now expanded for our rule-based system also to extract. RESULTS: This paper presents new results for extraction of relational information from biomedical text. We have improved our system, STRING-IE, to capture both new types of linguistic constructs as well as new types of biological information [i.e. (de-)phosphorylation]. The precision remains stable with a slight increase in recall. From almost one million PubMed abstracts related to four model organisms, we manage to extract regulatory networks and binary phosphorylations comprising 3,319 relation chunks. The accuracy is 83-90% and 86-95% for gene expression and (de-)phosphorylation relations, respectively. To achieve this, we made use of an organism-specific resource of gene/protein names considerably larger than those used in most other biology related information extraction approaches. These names were included in the lexicon when retraining the part-of-speech (POS) tagger on the GENIA corpus. For the domain in question, an accuracy of 96.4% was attained on POS tags. It should be noted that the rules were developed for yeast and successfully applied to both abstracts and full-text articles related to other organisms with comparable accuracy. AVAILABILITY: The revised GENIA corpus, the POS tagger, the extraction rules and the full sets of extracted relations are available from http://www.bork.embl.de/Docu/STRING-IE

Abstracting and Indexing↗

DCD - a novel plant specific domain in proteins involved in development and programmed cell death.

BACKGROUND: Recognition of microbial pathogens by plants triggers the hypersensitive reaction, a common form of programmed cell death in plants. These dying cells generate signals that activate the plant immune system and alarm the neighboring cells as well as the whole plant to activate defense responses to limit the spread of the pathogen. The molecular mechanisms behind the hypersensitive reaction are largely unknown except for the recognition process of pathogens. We delineate the NRP-gene in soybean, which is specifically induced during this programmed cell death and contains a novel protein domain, which is commonly found in different plant proteins. RESULTS: The sequence analysis of the protein, encoded by the NRP-gene from soybean, led to the identification of a novel domain, which we named DCD, because it is found in plant proteins involved in development and cell death. The domain is shared by several proteins in the Arabidopsis and the rice genomes, which otherwise show a different protein architecture. Biological studies indicate a role of these proteins in phytohormone response, embryo development and programmed cell by pathogens or ozone. CONCLUSION: It is tempting to speculate, that the DCD domain mediates signaling in plant development and programmed cell death and could thus be used to identify interacting proteins to gain further molecular insights into these processes.

Amino Acid Motifs↗

Structural genomics of human proteins--target selection and generation of a public catalogue of expression clones.

BACKGROUND: The availability of suitable recombinant protein is still a major bottleneck in protein structure analysis. The Protein Structure Factory, part of the international structural genomics initiative, targets human proteins for structure determination. It has implemented high throughput procedures for all steps from cloning to structure calculation. This article describes the selection of human target proteins for structure analysis, our high throughput cloning strategy, and the expression of human proteins in Escherichia coli host cells. RESULTS AND CONCLUSION: Protein expression and sequence data of 1414 E. coli expression clones representing 537 different proteins are presented. 139 human proteins (18%) could be expressed and purified in soluble form and with the expected size. All E. coli expression clones are publicly available to facilitate further functional characterisation of this set of human proteins.

Journal Article↗

Extraction of transcript diversity from scientific literature.

Transcript diversity generated by alternative splicing and associated mechanisms contributes heavily to the functional complexity of biological systems. The numerous examples of the mechanisms and functional implications of these events are scattered throughout the scientific literature. Thus, it is crucial to have a tool that can automatically extract the relevant facts and collect them in a knowledge base that can aid the interpretation of data from high-throughput methods. We have developed and applied a composite text-mining method for extracting information on transcript diversity from the entire MEDLINE database in order to create a database of genes with alternative transcripts. It contains information on tissue specificity, number of isoforms, causative mechanisms, functional implications, and experimental methods used for detection. We have mined this resource to identify 959 instances of tissue-specific splicing. Our results in combination with those from EST-based methods suggest that alternative splicing is the preferred mechanism for generating transcript diversity in the nervous system. We provide new annotations for 1,860 genes with the potential for generating transcript diversity. We assign the MeSH term "alternative splicing" to 1,536 additional abstracts in the MEDLINE database and suggest new MeSH terms for other events. We have successfully extracted information about transcript diversity and semiautomatically generated a database, LSAT, that can provide a quantitative understanding of the mechanisms behind tissue-specific gene expression. LSAT (Literature Support for Alternative Transcripts) is publicly available at http://www.bork.embl.de/LSAT/.

Journal Article↗

Comparative metagenomics of microbial communities.

The species complexity of microbial communities and challenges in culturing representative isolates make it difficult to obtain assembled genomes. Here we characterize and compare the metabolic capabilities of terrestrial and marine microbial communities using largely unassembled sequence data obtained by shotgun sequencing DNA isolated from the various environments. Quantitative gene content analysis reveals habitat-specific fingerprints that reflect known characteristics of the sampled environments. The identification of environment-specific genes through a gene-centric comparative analysis presents new opportunities for interpreting and diagnosing environments.

Animals↗