PubMed Health⌕ Search

Biomedical subjects

Peer Bork

Publications and source records attributed to Peer Bork.

At least 19 recordsLinked to original sources

microntology: a lightweight, data-driven controlled vocabulary to describe earth's microbial habitats.

MOTIVATION: Data-enabled studies of microbial ecology and evolution depend on high-quality descriptions of microbial habitats, based on curated and consolidated vocabularies. RESULTS: We introduce microntology v1.0, a pragmatic controlled vocabulary of 148 terms to describe microbial habitats and lifestyles, and provide manually curated microntology annotations for >300k metagenomic samples from public repositories. AVAILABILITY: microntology controlled vocabulary terms and term hierarchies (doi: 10.5281/zenodo.19730167), and curated annotations for 305 626 metagenomic samples (doi: 10.5281/zenodo.18164252) are available via Zenodo and spire.embl.de/downloads. Underlying code is available via github.com/grp-schmidt/microntology and Zenodo (doi: 10.5281/zenodo.20323497). User feedback, suggestions and bug reports are welcome at github.com/grp-schmidt/microntology/issues.

Ecosystem↗

Nonsense-mediated mRNA decay in Drosophila: at the intersection of the yeast and mammalian pathways.

The nonsense-mediated mRNA decay (NMD) pathway promotes the rapid degradation of mRNAs containing premature stop codons (PTCs). In Caenorhabditis elegans, seven genes (smg1-7) playing an essential role in NMD have been identified. Only SMG2-4 (known as UPF1-3) have orthologs in Saccharomyces cerevisiae. Here we show that the Drosophila orthologs of UPF1-3, SMG1, SMG5 and SMG6 are required for the degradation of PTC-containing mRNAs, but that there is no SMG7 ortholog in this organism. In contrast, orthologs of SMG5-7 are encoded by the human genome and all three are required for NMD. In human cells, exon boundaries have been shown to play a critical role in defining PTCs. This role is mediated by components of the exon junction complex (EJC). Contrary to expectation, however, we show that the components of the EJC are dispensable for NMD in Drosophila cells. Consistently, PTC definition occurs independently of exon boundaries in Drosophila. Our findings reveal that despite conservation of the NMD machinery, different mechanisms have evolved to discriminate premature from natural stop codons in metazoa.

Amino Acid Sequence↗

The DNA sequence of human chromosome 7.

Human chromosome 7 has historically received prominent attention in the human genetics community, primarily related to the search for the cystic fibrosis gene and the frequent cytogenetic changes associated with various forms of cancer. Here we present more than 153 million base pairs representing 99.4% of the euchromatic sequence of chromosome 7, the first metacentric chromosome completed so far. The sequence has excellent concordance with previously established physical and genetic maps, and it exhibits an unusual amount of segmentally duplicated sequence (8.2%), with marked differences between the two arms. Our initial analyses have identified 1,150 protein-coding genes, 605 of which have been confirmed by complementary DNA sequences, and an additional 941 pseudogenes. Of genes confirmed by transcript sequences, some are polymorphic for mutations that disrupt the reading frame.

Animals↗

Update on XplorMed: A web server for exploring scientific literature.

As scientific literature databases like MEDLINE increase in size, so does the time required to search them. Scientists must frequently inspect long lists of references manually, often just reading the titles. XplorMed is a web tool that aids MEDLINE searching by summarizing the subjects contained in the results, thus allowing users to focus on subjects of interest. Here we describe new features added to XplorMed during the last 2 years (http://www.bork.embl-heidelberg.de/xplormed/).

Bibliography of Medicine↗

ELM server: A new resource for investigating short functional sites in modular eukaryotic proteins.

Multidomain proteins predominate in eukaryotic proteomes. Individual functions assigned to different sequence segments combine to create a complex function for the whole protein. While on-line resources are available for revealing globular domains in sequences, there has hitherto been no comprehensive collection of small functional sites/motifs comparable to the globular domain resources, yet these are as important for the function of multidomain proteins. Short linear peptide motifs are used for cell compartment targeting, protein-protein interaction, regulation by phosphorylation, acetylation, glycosylation and a host of other post-translational modifications. ELM, the Eukaryotic Linear Motif server at http://elm.eu.org/, is a new bioinformatics resource for investigating candidate short non-globular functional motifs in eukaryotic proteins, aiming to fill the void in bioinformatics tools. Sequence comparisons with short motifs are difficult to evaluate because the usual significance assessments are inappropriate. Therefore the server is implemented with several logical filters to eliminate false positives. Current filters are for cell compartment, globular domain clash and taxonomic range. In favourable cases, the filters can reduce the number of retained matches by an order of magnitude or more.

Amino Acid Motifs↗

Systematic discovery of analogous enzymes in thiamin biosynthesis.

In all genome-sequencing projects completed to date, a considerable number of 'gaps' have been found in the biochemical pathways of the respective species. In many instances, missing enzymes are displaced by analogs, functionally equivalent proteins that have evolved independently and lack sequence and structural similarity. Here we fill such gaps by analyzing anticorrelating occurrences of genes across species. Our approach, applied to the thiamin biosynthesis pathway comprising approximately 15 catalytic steps, predicts seven instances in which known enzymes have been displaced by analogous proteins. So far we have verified four predictions by genetic complementation, including three proteins for which there was no previous experimental evidence of a role in the thiamin biosynthesis pathway. For one hypothetical protein, biochemical characterization confirmed the predicted thiamin phosphate synthase (ThiE) activity. The results demonstrate the ability of our computational approach to predict specific functions without taking into account sequence similarity.

Alkyl and Aryl Transferases↗

Information extraction from full text scientific articles: where are the keywords?

BACKGROUND: To date, many of the methods for information extraction of biological information from scientific articles are restricted to the abstract of the article. However, full text articles in electronic version, which offer larger sources of data, are currently available. Several questions arise as to whether the effort of scanning full text articles is worthy, or whether the information that can be extracted from the different sections of an article can be relevant. RESULTS: In this work we addressed those questions showing that the keyword content of the different sections of a standard scientific article (abstract, introduction, methods, results, and discussion) is very heterogeneous. CONCLUSIONS: Although the abstract contains the best ratio of keywords per total of words, other sections of the article may be a better source of biologically relevant data.

Anatomy↗

Pathogenesis of DNA repair-deficient cancers: a statistical meta-analysis of putative Real Common Target genes.

DNA mismatch repair deficiency is observed in about 15% of human colorectal, gastric, and endometrial tumors and in lower frequencies in a minority of other tumors thereby causing insertion/deletion mutations at short repetitive sequences, recognized as microsatellite instability (MSI). Evolution of tumors, including those with MSI, is a continuous process of mutation and selection favoring neoplastic growth. Mutations in microsatellite-bearing genes that promote tumor cell growth in general (Real Common Target genes) are assumed to be the driving force during MSI carcinogenesis. Thus, microsatellite mutations in these genes should occur more frequently than mutations in microsatellite genes without contribution to malignancy (ByStander genes). So far, only a few Real Common Target genes have been identified by functional studies. Thus, comprehensive analysis of microsatellite mutations will provide important clues to the understanding of MSI-driven carcinogenesis. Here, we evaluated published mutation frequencies on 194 repeat tracts in 137 genes in MSI-H colorectal, endometrial, and gastric carcinomas and propose a statistical model that aims to identify Real Common Target genes. According to our model nine genes including BAX and TGFbetaRII were identified as Real Common Targets in colorectal cancer, one gene in gastric cancer, and three genes in endometrial cancer. Microsatellite mutations in five additional genes seem to be counterselected in gastrointestinal tumors. Overall, the general applicability, the capacity to unlimited data analysis, the inclusion of mutation data generated by different groups on different sets of tumors make this model a useful tool for predicting Real Common Target genes with specificity for MSI-H tumors of different organs, guiding subsequent functional studies to the most likely targets among numerous microsatellite harboring genes.

Base Pair Mismatch↗

STRING: a database of predicted functional associations between proteins.

Functional links between proteins can often be inferred from genomic associations between the genes that encode them: groups of genes that are required for the same function tend to show similar species coverage, are often located in close proximity on the genome (in prokaryotes), and tend to be involved in gene-fusion events. The database STRING is a precomputed global resource for the exploration and analysis of these associations. Since the three types of evidence differ conceptually, and the number of predicted interactions is very large, it is essential to be able to assess and compare the significance of individual predictions. Thus, STRING contains a unique scoring-framework based on benchmarks of the different types of associations against a common reference set, integrated in a single confidence score per prediction. The graphical representation of the network of inferred, weighted protein interactions provides a high-level view of functional linkage, facilitating the analysis of modularity in biological processes. STRING is updated continuously, and currently contains 261 033 orthologs in 89 fully sequenced genomes. The database predicts functional interactions at an expected level of accuracy of at least 80% for more than half of the genes; it is online at http://www.bork.embl-heidelberg.de/STRING/.

Algorithms↗

The InterPro Database, 2003 brings increased coverage and new features.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created in 1999 as a means of amalgamating the major protein signature databases into one comprehensive resource. PROSITE, Pfam, PRINTS, ProDom, SMART and TIGRFAMs have been manually integrated and curated and are available in InterPro for text- and sequence-based searching. The results are provided in a single format that rationalises the results that would be obtained by searching the member databases individually. The latest release of InterPro contains 5629 entries describing 4280 families, 1239 domains, 95 repeats and 15 post-translational modifications. Currently, the combined signatures in InterPro cover more than 74% of all proteins in SWISS-PROT and TrEMBL, an increase of nearly 15% since the inception of InterPro. New features of the database include improved searching capabilities and enhanced graphical user interfaces for visualisation of the data. The database is available via a webserver (http://www.ebi.ac.uk/interpro) and anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Animals↗

Increase of functional diversity by alternative splicing.

A large-scale analysis of protein isoforms arising from alternative splicing shows that alternative splicing tends to insert or delete complete protein domains more frequently than expected by chance, whereas disruption of domains and other structural modules is less frequent. If domain regions are disrupted, the functional effect, as predicted from 3D structure, is frequently equivalent to removal of the entire domain. Also, short alternative splicing events within domains, which might preserve folded structure, target functional residues more frequently than expected. Thus, it seems that positive selection has had a major role in the evolution of alternative splicing.

Alternative Splicing↗

The identification of a conserved domain in both spartin and spastin, mutated in hereditary spastic paraplegia.

Multiple sequence alignment has revealed the presence of a sequence domain of approximately 80 amino acids in two molecules, spartin and spastin, mutated in hereditary spastic paraplegia. The domain, which corresponds to a slightly extended version of the recently described ESP domain of unknown function, was also identified in VPS4, SKD1, RPK118, and SNX15, all of which have a well established and consistent role in endosomal trafficking. Recent functional information indicates that spastin is likely to be involved in microtubule interaction. With this new information relating to its likely function, we propose the more descriptive name 'MIT' (contained within microtubule-interacting and trafficking molecules) for the domain and predict endosomal trafficking as the principal functionality of all molecules in which it is present.

Adenosine Triphosphatases↗

Function prediction and protein networks.

In the genomics era, the interactions between proteins are at the center of attention. Genomic-context methods used to predict these interactions have been put on a quantitative basis, revealing that they are at least on an equal footing with genomics experimental data. A survey of experimentally confirmed predictions proves the applicability of these methods, and new concepts to predict protein interactions in eukaryotes have been described. Finally, the interaction networks that can be obtained by combining the predicted pair-wise interactions have enough internal structure to detect higher levels of organization, such as 'functional modules'.

Animals↗

Metabolites: a helping hand for pathway evolution?

The evolution of enzymes and pathways is under debate. Recent studies show that recruitment of single enzymes from different pathways could be the driving force for pathway evolution. Other mechanisms of evolution, such as pathway duplication, enzyme specialization, de novo invention of pathways or retro-evolution of pathways, appear to be less abundant. Twenty percent of enzyme superfamilies are quite variable, not only in changing reaction chemistry or metabolite type but in changing both at the same time. These variable superfamilies account for nearly half of all known reactions. The most frequently occurring metabolites provide a helping hand for such changes because they can be accommodated by many enzyme superfamilies. Thus, a picture is emerging in which new pathways are evolving from central metabolites by preference, thereby keeping the overall topology of the metabolic network.

Animals↗

Bioinformatics in the post-sequence era.

In the past decade, bioinformatics has become an integral part of research and development in the biomedical sciences. Bioinformatics now has an essential role both in deciphering genomic, transcriptomic and proteomic data generated by high-throughput experimental technologies and in organizing information gathered from traditional biology. Sequence-based methods of analyzing individual genes or proteins have been elaborated and expanded, and methods have been developed for analyzing large numbers of genes or proteins simultaneously, such as in the identification of clusters of related genes and networks of interacting proteins. With the complete genome sequences for an increasing number of organisms at hand, bioinformatics is beginning to provide both conceptual bases and practical methods for detecting systemic functional behaviors of the cell and the organism.

Computational Biology↗

The way we write.

Explore the source record for details and available documents.

Biomedical Research↗