PubMed Health⌕ Search

Biomedical subjects

Peer Bork

Publications and source records attributed to Peer Bork.

At least 91 records · Page 5Linked to original sources

Information extraction from full text scientific articles: where are the keywords?

BACKGROUND: To date, many of the methods for information extraction of biological information from scientific articles are restricted to the abstract of the article. However, full text articles in electronic version, which offer larger sources of data, are currently available. Several questions arise as to whether the effort of scanning full text articles is worthy, or whether the information that can be extracted from the different sections of an article can be relevant. RESULTS: In this work we addressed those questions showing that the keyword content of the different sections of a standard scientific article (abstract, introduction, methods, results, and discussion) is very heterogeneous. CONCLUSIONS: Although the abstract contains the best ratio of keywords per total of words, other sections of the article may be a better source of biologically relevant data.

Anatomy↗

Pathogenesis of DNA repair-deficient cancers: a statistical meta-analysis of putative Real Common Target genes.

DNA mismatch repair deficiency is observed in about 15% of human colorectal, gastric, and endometrial tumors and in lower frequencies in a minority of other tumors thereby causing insertion/deletion mutations at short repetitive sequences, recognized as microsatellite instability (MSI). Evolution of tumors, including those with MSI, is a continuous process of mutation and selection favoring neoplastic growth. Mutations in microsatellite-bearing genes that promote tumor cell growth in general (Real Common Target genes) are assumed to be the driving force during MSI carcinogenesis. Thus, microsatellite mutations in these genes should occur more frequently than mutations in microsatellite genes without contribution to malignancy (ByStander genes). So far, only a few Real Common Target genes have been identified by functional studies. Thus, comprehensive analysis of microsatellite mutations will provide important clues to the understanding of MSI-driven carcinogenesis. Here, we evaluated published mutation frequencies on 194 repeat tracts in 137 genes in MSI-H colorectal, endometrial, and gastric carcinomas and propose a statistical model that aims to identify Real Common Target genes. According to our model nine genes including BAX and TGFbetaRII were identified as Real Common Targets in colorectal cancer, one gene in gastric cancer, and three genes in endometrial cancer. Microsatellite mutations in five additional genes seem to be counterselected in gastrointestinal tumors. Overall, the general applicability, the capacity to unlimited data analysis, the inclusion of mutation data generated by different groups on different sets of tumors make this model a useful tool for predicting Real Common Target genes with specificity for MSI-H tumors of different organs, guiding subsequent functional studies to the most likely targets among numerous microsatellite harboring genes.

Base Pair Mismatch↗

STRING: a database of predicted functional associations between proteins.

Functional links between proteins can often be inferred from genomic associations between the genes that encode them: groups of genes that are required for the same function tend to show similar species coverage, are often located in close proximity on the genome (in prokaryotes), and tend to be involved in gene-fusion events. The database STRING is a precomputed global resource for the exploration and analysis of these associations. Since the three types of evidence differ conceptually, and the number of predicted interactions is very large, it is essential to be able to assess and compare the significance of individual predictions. Thus, STRING contains a unique scoring-framework based on benchmarks of the different types of associations against a common reference set, integrated in a single confidence score per prediction. The graphical representation of the network of inferred, weighted protein interactions provides a high-level view of functional linkage, facilitating the analysis of modularity in biological processes. STRING is updated continuously, and currently contains 261 033 orthologs in 89 fully sequenced genomes. The database predicts functional interactions at an expected level of accuracy of at least 80% for more than half of the genes; it is online at http://www.bork.embl-heidelberg.de/STRING/.

Algorithms↗

The InterPro Database, 2003 brings increased coverage and new features.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created in 1999 as a means of amalgamating the major protein signature databases into one comprehensive resource. PROSITE, Pfam, PRINTS, ProDom, SMART and TIGRFAMs have been manually integrated and curated and are available in InterPro for text- and sequence-based searching. The results are provided in a single format that rationalises the results that would be obtained by searching the member databases individually. The latest release of InterPro contains 5629 entries describing 4280 families, 1239 domains, 95 repeats and 15 post-translational modifications. Currently, the combined signatures in InterPro cover more than 74% of all proteins in SWISS-PROT and TrEMBL, an increase of nearly 15% since the inception of InterPro. New features of the database include improved searching capabilities and enhanced graphical user interfaces for visualisation of the data. The database is available via a webserver (http://www.ebi.ac.uk/interpro) and anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Animals↗

Alternative splicing and evolution.

Alternative splicing is a critical post-transcriptional event leading to an increase in the transcriptome diversity. Recent bioinformatics studies revealed a high frequency of alternative splicing. Although the extent of AS conservation among mammals is still being discussed, it has been argued that major forms of alternatively spliced transcripts are much better conserved than minor forms. It suggests that alternative splicing plays a major role in genome evolution allowing new exons to evolve with less constraint.

Alternative Splicing↗

Protein disorder prediction: implications for structural proteomics.

A great challenge in the proteomics and structural genomics era is to predict protein structure and function, including identification of those proteins that are partially or wholly unstructured. Disordered regions in proteins often contain short linear peptide motifs (e.g., SH3 ligands and targeting signals) that are important for protein function. We present here DisEMBL, a computational tool for prediction of disordered/unstructured regions within a protein sequence. As no clear definition of disorder exists, we have developed parameters based on several alternative definitions and introduced a new one based on the concept of "hot loops," i.e., coils with high temperature factors. Avoiding potentially disordered segments in protein expression constructs can increase expression, foldability, and stability of the expressed protein. DisEMBL is thus useful for target selection and the design of constructs as needed for many biochemical studies, particularly structural biology and structural genomics projects. The tool is freely available via a web interface (http://dis.embl.de) and can be downloaded for use in large-scale studies.

Circular Dichroism↗

Increase of functional diversity by alternative splicing.

A large-scale analysis of protein isoforms arising from alternative splicing shows that alternative splicing tends to insert or delete complete protein domains more frequently than expected by chance, whereas disruption of domains and other structural modules is less frequent. If domain regions are disrupted, the functional effect, as predicted from 3D structure, is frequently equivalent to removal of the entire domain. Also, short alternative splicing events within domains, which might preserve folded structure, target functional residues more frequently than expected. Thus, it seems that positive selection has had a major role in the evolution of alternative splicing.

Alternative Splicing↗

The identification of a conserved domain in both spartin and spastin, mutated in hereditary spastic paraplegia.

Multiple sequence alignment has revealed the presence of a sequence domain of approximately 80 amino acids in two molecules, spartin and spastin, mutated in hereditary spastic paraplegia. The domain, which corresponds to a slightly extended version of the recently described ESP domain of unknown function, was also identified in VPS4, SKD1, RPK118, and SNX15, all of which have a well established and consistent role in endosomal trafficking. Recent functional information indicates that spastin is likely to be involved in microtubule interaction. With this new information relating to its likely function, we propose the more descriptive name 'MIT' (contained within microtubule-interacting and trafficking molecules) for the domain and predict endosomal trafficking as the principal functionality of all molecules in which it is present.

Adenosine Triphosphatases↗

Function prediction and protein networks.

In the genomics era, the interactions between proteins are at the center of attention. Genomic-context methods used to predict these interactions have been put on a quantitative basis, revealing that they are at least on an equal footing with genomics experimental data. A survey of experimentally confirmed predictions proves the applicability of these methods, and new concepts to predict protein interactions in eukaryotes have been described. Finally, the interaction networks that can be obtained by combining the predicted pair-wise interactions have enough internal structure to detect higher levels of organization, such as 'functional modules'.

Animals↗

Metabolites: a helping hand for pathway evolution?

The evolution of enzymes and pathways is under debate. Recent studies show that recruitment of single enzymes from different pathways could be the driving force for pathway evolution. Other mechanisms of evolution, such as pathway duplication, enzyme specialization, de novo invention of pathways or retro-evolution of pathways, appear to be less abundant. Twenty percent of enzyme superfamilies are quite variable, not only in changing reaction chemistry or metabolite type but in changing both at the same time. These variable superfamilies account for nearly half of all known reactions. The most frequently occurring metabolites provide a helping hand for such changes because they can be accommodated by many enzyme superfamilies. Thus, a picture is emerging in which new pathways are evolving from central metabolites by preference, thereby keeping the overall topology of the metabolic network.

Animals↗

Bioinformatics in the post-sequence era.

In the past decade, bioinformatics has become an integral part of research and development in the biomedical sciences. Bioinformatics now has an essential role both in deciphering genomic, transcriptomic and proteomic data generated by high-throughput experimental technologies and in organizing information gathered from traditional biology. Sequence-based methods of analyzing individual genes or proteins have been elaborated and expanded, and methods have been developed for analyzing large numbers of genes or proteins simultaneously, such as in the identification of clusters of related genes and networks of interacting proteins. With the complete genome sequences for an increasing number of organisms at hand, bioinformatics is beginning to provide both conceptual bases and practical methods for detecting systemic functional behaviors of the cell and the organism.

Computational Biology↗

The way we write.

Explore the source record for details and available documents.

Biomedical Research↗

A genome-wide survey of human pseudogenes.

We screened all intergenic regions in the human genome to identify pseudogenes with a combination of homology searches and a functionality test using the ratio of silent to replacement nucleotide substitutions (KA/KS). We identified 19,724 regions of which 95% +/- 3% are estimated to evolve neutrally and thus are likely to encode pseudogenes. Half of these have no detectable truncation in their pseudocoding regions and therefore are not identifiable by methods that require the presence of truncations to prove nonfunctionality. A comparative analysis with the mouse genome showed that 70% of these pseudogenes have a retrotranspositional origin (processed), and the rest arose by segmental duplication (nonprocessed). Although the spread of both types of pseudogenes correlates with chromosome size, nonprocessed pseudogenes appear to be enriched in regions with high gene density. It is likely that the human pseudogenes identified here represent only a small fraction of the total, which probably exceeds the number of genes.

Benchmarking↗

A protocol for the update of references to scientific literature in biological databases.

Entries in biological databases are usually linked to scientific references. To generate those links and to keep them up-to-date, database maintainers have to continuously scan the scientific literature to select references that are relevant for each single database entry. The continuous growth of both the corpus of scientific literature and the size of biological databases makes this task very hard. We present a protocol intended to assist the updating of an existing set of literature (abstract) links from a single database entry with new references. It consists of taking the set of MEDLINE neighbour references of the existing linked abstracts and evaluating their relevance according to the existing set of abstracts. To test the applicability of the algorithm, we did a simple benchmark of the system using the references associated with the entries of a protein domain database. Human experts found the references that the algorithm scored highly were more relevant to the database entry than those scored lowly, suggesting that the algorithm was useful.

Abstracting and Indexing↗

Identification and characterization of UEV3, a human cDNA with similarities to inactive E2 ubiquitin-conjugating enzymes.

Recent studies have shown that ubiquitination is an essential factor in endosomal sorting and virus assembly. The human TSG101 gene has been demonstrated to belong to a group of genes coding for apparently inactive E2 ubiquitin-conjugating enzymes, which exert regulatory effects on E2 activity in cellular ubiquitination processes. In this study, a novel human cDNA (UEV3) encoding a putative protein of 379 amino acids was isolated from a human placenta library that may represent a partial paralogue of human TSG101. The predicted protein contains an N-terminal domain homologous to the catalytic domain of ubiquitin-conjugating enzymes (Ubc), which is fused to a sequence showing significant homology to members of the lactate dehydrogenase protein family. The UEV3 gene is located on chromosome 11 closely adjacent to TSG101 and LDH-C. Northern blot and UEV3-specific reverse transcription/polymerase chain reaction (RT/PCR) analyses of various colon carcinoma cell lines as well as both normal and tumor samples from colon revealed an expression of the UEV3 cDNA in all tested samples.

Amino Acid Sequence↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗