PubMed Health⌕ Search

Biomedical subjects

Peer Bork

Publications and source records attributed to Peer Bork.

123 records · Page 7Linked to original sources

CASH--a beta-helix domain widespread among carbohydrate-binding proteins.

In this article, we describe a novel, widespread domain (CASH) that is shared by many carbohydrate-binding proteins and sugar hydrolases. This domain occurs in more than 1000 proteins distributed among all three kingdoms of life. The CASH domain is characterized by internal repetitions of glycines and hydrophobic residues that correspond to the repetitive units of a predicted or observed right-handed beta-helix structure of the pectate lyase superfamily.

Amino Acid Motifs↗

AMOP, a protein module alternatively spliced in cancer cells.

This article describes a new extracellular domain--AMOP, for adhesion-associated domain in MUC4 and other proteins. This domain occurs in putative cell adhesion molecules and in some splice variants of MUC4. MUC4 splice variants are overexpressed in several tumours; in particular, they are highly expressed in pancreatic carcinomas but not in normal pancreas. The presence of AMOP in cell adhesion molecules could be indicative of a role for this domain in adhesion.

Alternative Splicing↗

Genome and protein evolution in eukaryotes.

The past year has seen the completion of the genome sequence of the flowering plant Arabidopsis thaliana and the initial sequence reports of the human genome. The availability of completely sequenced eukaryotic genomes from disparate phylogenetic lineages has opened the door to comparative analyses and a better understanding of the evolutionary processes shaping genomes. Complex many-to-many relationships between genes from different species appear to be the norm, suggesting that transfer of detailed functional annotation will not be straightforward. In addition to expansion and contraction of gene families, new genes evolve from recombination of pre-existing domains, although some domain families do appear to have evolved recently and to be specific to restricted phylogenetic lineages. The overall picture is of a huge diversity of gene content within eukaryotic genomes, reflecting different functional demands in different species.

Animals↗

InterPro: an integrated documentation resource for protein families, domains and functional sites.

The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.

Algorithms↗

Comparative analysis of protein interaction networks.

Recent advances in proteomics and computational biology have lead to a flood of protein interaction data and resulting interaction networks (e.g. (Gavin et al., 2002)). Here I first analyse the status and quality of parts lists (genes and proteins), then comparatively assess large-scale protein interaction data (von Mering et al., 2002) and finally try to identify biological meaningful units (e.g. pathways, cellular processes) within interaction networks that are derived from the conservation of gene neighborhood (Snel et al., 2002). Possible extensions of gene neighborhood analysis to eukaryotes (von Mering and Bork, 2002) will be discussed.

Animals↗

A complex prediction: three-dimensional model of the yeast exosome.

We present a model of the yeast exosome based on the bacterial degradosome component polynucleotide phosphorylase (PNPase). Electron microscopy shows the exosome to resemble PNPase but with key differences likely related to the position of RNA binding domains, and to the location of domains unique to the exosome. We use various techniques to reduce the many possible models of exosome subunits based on PNPase to just one. The model suggests numerous experiments to probe exosome function, particularly with respect to subunits making direct atomic contacts and conserved, possibly functional residues within the predicted central pore of the complex.

Amino Acid Sequence↗

The rhodanese/Cdc25 phosphatase superfamily. Sequence-structure-function relations.

Rhodanese domains are ubiquitous structural modules occurring in the three major evolutionary phyla. They are found as tandem repeats, with the C-terminal domain hosting the properly structured active-site Cys residue, as single domain proteins or in combination with distinct protein domains. An increasing number of reports indicate that rhodanese modules are versatile sulfur carriers that have adapted their function to fulfill the need for reactive sulfane sulfur in distinct metabolic and regulatory pathways. Recent investigations have shown that rhodanese domains are also structurally related to the catalytic subunit of Cdc25 phosphatase enzymes and that the two enzyme families are likely to share a common evolutionary origin. In this review, the rhodanese/Cdc25 phosphatase superfamily is analyzed. Although the identification of their biological substrates has thus far proven elusive, the emerging picture points to a role for the amino-acid composition of the active-site loop in substrate recognition/specificity. Furthermore, the frequently observed association of catalytically inactive rhodanese modules with other protein domains suggests a distinct regulatory role for these inactive domains, possibly in connection with signaling.

Animals↗

Genomes in flux: the evolution of archaeal and proteobacterial gene content.

In the course of evolution, genomes are shaped by processes like gene loss, gene duplication, horizontal gene transfer, and gene genesis (the de novo origin of genes). Here we reconstruct the gene content of ancestral Archaea and Proteobacteria and quantify the processes connecting them to their present day representatives based on the distribution of genes in completely sequenced genomes. We estimate that the ancestor of the Proteobacteria contained around 2500 genes, and the ancestor of the Archaea around 2050 genes. Although it is necessary to invoke horizontal gene transfer to explain the content of present day genomes, gene loss, gene genesis, and simple vertical inheritance are quantitatively the most dominant processes in shaping the genome. Together they result in a turnover of gene content such that even the lineage leading from the ancestor of the Proteobacteria to the relatively large genome of Escherichia coli has lost at least 950 genes. Gene loss, unlike the other processes, correlates fairly well with time. This clock-like behavior suggests that gene loss is under negative selection, while the processes that add genes are under positive selection.

Amino Acid Substitution↗

Systematic identification of novel protein domain families associated with nuclear functions.

A systematic computational analysis of protein sequences containing known nuclear domains led to the identification of 28 novel domain families. This represents a 26% increase in the starting set of 107 known nuclear domain families used for the analysis. Most of the novel domains are present in all major eukaryotic lineages, but 3 are species specific. For about 500 of the 1200 proteins that contain these new domains, nuclear localization could be inferred, and for 700, additional features could be predicted. For example, we identified a new domain, likely to have a role downstream of the unfolded protein response; a nematode-specific signalling domain; and a widespread domain, likely to be a noncatalytic homolog of ubiquitin-conjugating enzymes.

Amidohydrolases↗

Predicting protein cellular localization using a domain projection method.

We investigate the co-occurrence of domain families in eukaryotic proteins to predict protein cellular localization. Approximately half (300) of SMART domains form a "small-world network", linked by no more than seven degrees of separation. Projection of the domains onto two-dimensional space reveals three clusters that correspond to cellular compartments containing secreted, cytoplasmic, and nuclear proteins. The projection method takes into account the existence of "bridging" domains, that is, instances where two domains might not occur with each other but frequently co-occur with a third domain; in such circumstances the domains are neighbors in the projection. While the majority of domains are specific to a compartment ("locale"), and hence may be used to localize any protein that contains such a domain, a small subset of domains either are present in multiple locales or occur in transmembrane proteins. Comparison with previously annotated proteins shows that SMART domain data used with this approach can predict, with 92% accuracy, the localizations of 23% of eukaryotic proteins. The coverage and accuracy will increase with improvements in domain database coverage. This method is complementary to approaches that use amino-acid composition or identify sorting sequences; these methods may be combined to further enhance prediction accuracy.

Animals↗

Exploring MEDLINE abstracts with XplorMed.

XplorMed is a publicly available web tool conceived to make life easier for MEDLINE(c) users looking for scientific information. Searching scientific literature is an information retrieval problem. Abstracts that are of possible interest to the user are usually selected by a keyword search followed by manual screening, which often results in the retrieval of a large number of abstracts. Interesting references can be buried among irrelevant ones because of nonspecific queries. XplorMed is intended to extract dependency relations between the words of the abstracts. These relations can be filtered and arranged to deduce different subjects in the query and offer a condensed view of the abstract, allowing users to select texts of interest without having to read them all. XplorMed is available http://www.bork. embl-heidelberg.de/xplormed.

Artificial Intelligence↗

Computing fuzzy associations for the analysis of biological literature.

The increase of information in biology makes it difficult for researchers in any field to keep current with the literature. The MEDLINE database of scientific abstracts can be quickly scanned using electronic mechanisms. Potentially interesting abstracts can be selected by matching words joined by Boolean operators. However this means of selecting documents is not optimal. Nonspecific queries have to be effected, resulting in large numbers of irrelevant abstracts that have to be manually scanned To facilitate this analysis, we have developed a system that compiles a summary of subjects and related documents on the results of a MEDLINE query. For this, we have applied a fuzzy binary relation formalism that deduces relations between words present in a set of abstracts preprocessed with a standard grammatical tagger. Those relations are used to derive ensembles of related words and their associated subsets of abstracts. The algorithm can be used publicly at http:// www.bork.embl-heidelberg.de/xplormed/.

Algorithms↗

Alternative splicing and genome complexity.

Alternative splicing of mRNA allows many gene products with different functions to be produced from a single coding sequence. It has recently been proposed as a mechanism by which higher-order diversity is generated. Here we show, using large-scale expressed sequence tag (EST) analysis, that among seven different eukaryotes the amount of alternative splicing is comparable, with no large differences between humans and other animals.

Alternative Splicing↗

Genome speak.

Explore the source record for details and available documents.

Journal Article↗

A versatile structural domain analysis server using profile weight matrices.

The WEB tool "AnDom" assigns to a given protein sequence all experimentally determined structural domains contained within it, including multidomain and large proteins. The server uses profile specific matrices from custom generated multiple sequence alignments of all known SCOP domains (SCOP version 1.50). Prediction time is short allowing numerous applications for structural genomics including investigation of complex eucaryotic protein families. The WWW server is at http://www.bork.embl-heidelberg.de/AnDom, and profiles can be downloaded at ftp.bork.embl-heidelberg.de/pub/users/ schmidt/AnDom.

Amino Acid Sequence↗