PubMed Health⌕ Search

Biomedical subjects

Ulf Leser

Publications and source records attributed to Ulf Leser.

9 recordsLinked to original sources

Implications for molecular mechanisms of glycoprotein hormone receptors using a new sequence-structure-function analysis resource.

Comparison between wild-type and mutated glycoprotein hormone receptors (GPHRs), TSH receptor, FSH receptor, and LH-chorionic gonadotropin receptor is established to identify determinants involved in molecular activation mechanism. The basic aims of the current work are 1) the discrimination of receptor phenotypes according to the differences between activity states they represent, 2) the assignment of classified phenotypes to three-dimensional structural positions to reveal 3) functional-structural hot spots and 4) interrelations between determinants that are responsible for corresponding activity states. Because it is hard to survey the vast amount of pathogenic and site-directed mutations at GPHRs and to improve an almost isolated consideration of individual point mutations, we present a system for systematic and diversified sequence-structure-function analysis (http://www.fmp-berlin.de/ssfa). To combine all mutagenesis data into one set, we converted the functional data into unified scaled values. This at least enables their comparison in a rough classification manner. In this study we describe the compiled data set and a wide spectrum of functions for user-driven searches and classification of receptor functionalities such as cell surface expression, maximum of hormone binding capability, and basal as well as hormone-induced Galphas/Galphaq mediated cAMP/inositol phosphate accumulation. Complementary to known databases, our data set and bioinformatics tools allow functional and biochemical specificities to be linked with spatial features to reveal concealed structure-function relationships by a semiquantitative analysis. A comprehensive discrimination of specificities of pathogenic mutations and in vitro mutant phenotypes and their relation to signaling mechanisms of GPHRs demonstrates the utility of sequence-structure-function analysis. Moreover, new interrelations of determinants important for selective G protein-mediated activation of GPHRs are resumed.

Animals↗

AliBaba: PubMed as a graph.

UNLABELLED: The biomedical literature contains a wealth of information on associations between many different types of objects, such as protein-protein interactions, gene-disease associations and subcellular locations of proteins. When searching such information using conventional search engines, e.g. PubMed, users see the data only one-abstract at a time and 'hidden' in natural language text. AliBaba is an interactive tool for graphical summarization of search results. It parses the set of abstracts that fit a PubMed query and presents extracted information on biomedical objects and their relationships as a graphical network. AliBaba extracts associations between cells, diseases, drugs, proteins, species and tissues. Several filter options allow for a more focused search. Thus, researchers can grasp complex networks described in various articles at a glance. AVAILABILITY: http://alibaba.informatik.hu-berlin.de/

Abstracting and Indexing↗

How well are protein structures annotated in secondary databases?

We investigated to what extent Protein Data Bank (PDB) entries are annotated with second-party information based on existing cross-references between PDB and 15 other databases. We report 2 interesting findings. First, there is a clear "annotation gap" for structures less than 7 years old for secondary databases that are manually curated. Second, the examined databases overlap with each other quite well, dividing the PDB into 2 well-annotated thirds and one poorly annotated third. Both observations should be taken into account in any study depending on the selection of protein structures by their annotation.

Amino Acid Sequence↗

A query language for biological networks.

MOTIVATION: Many areas of modern biology are concerned with the management, storage, visualization, comparison and analysis of networks, but no appropriate query language for such complex data structures yet exists. RESULTS: We have designed and implemented the pathway query language (PQL) for querying large protein interaction or pathway databases. PQL is based on a simple graph data model with extensions reflecting properties of biological objects. Queries match subgraphs in the database based on node properties and paths between nodes. The syntax is easy to learn for anybody familiar with SQL. As an important feature, a query may require a certain structure in the database to exist but return a different subgraph. We have tested PQL queries on networks of up to 16,000 nodes and found it to scale very well. AVAILABILITY: The code is available on request from the author.

Computational Biology↗

Systematic feature evaluation for gene name recognition.

In task 1A of the BioCreAtIvE evaluation, systems had to be devised that recognize words and phrases forming gene or protein names in natural language sentences. We approach this problem by building a word classification system based on a sliding window approach with a Support Vector Machine, combined with a pattern-based post-processing for the recognition of phrases. The performance of such a system crucially depends on the type of features chosen for consideration by the classification method, such as pre- or postfixes, character n-grams, patterns of capitalization, or classification of preceding or following words. We present a systematic approach to evaluate the performance of different feature sets based on recursive feature elimination, RFE. Based on a systematic reduction of the number of features used by the system, we can quantify the impact of different feature sets on the results of the word classification problem. This helps us to identify descriptive features, to learn about the structure of the problem, and to design systems that are faster and easier to understand. We observe that the SVM is robust to redundant features. RFE improves the performance by 0.7%, compared to using the complete set of attributes. Moreover, a performance that is only 2.3% below this maximum can be obtained using fewer than 5% of the features.

Computational Biology↗

GandrKB--ontological microarray annotation and visualization.

SUMMARY: The Gandr (gene annotation data representation) knowledgebase is an ontological framework for laboratory-specific gene annotation. Gandr uses Protege 2000 for editing, querying and visualizing microarray data and annotations. Genes can be annotated with provided, newly created or imported ontological concepts. Annotated genes can inherit assigned concept properties and can be related to each other. The resulting knowledgebase can be visualized as interactive network of nodes and edges representing genes and their functional relationships. This allows for immediate and associative gene context exploration. Ontological query techniques allow for powerful data access.

Algorithms↗

Columba: an integrated database of proteins, structures, and annotations.

BACKGROUND: Structural and functional research often requires the computation of sets of protein structures based on certain properties of the proteins, such as sequence features, fold classification, or functional annotation. Compiling such sets using current web resources is tedious because the necessary data are spread over many different databases. To facilitate this task, we have created COLUMBA, an integrated database of annotations of protein structures. DESCRIPTION: COLUMBA currently integrates twelve different databases, including PDB, KEGG, Swiss-Prot, CATH, SCOP, the Gene Ontology, and ENZYME. The database can be searched using either keyword search or data source-specific web forms. Users can thus quickly select and download PDB entries that, for instance, participate in a particular pathway, are classified as containing a certain CATH architecture, are annotated as having a certain molecular function in the Gene Ontology, and whose structures have a resolution under a defined threshold. The results of queries are provided in both machine-readable extensible markup language and human-readable format. The structures themselves can be viewed interactively on the web. CONCLUSION: The COLUMBA database facilitates the creation of protein structure data sets for many structure-based studies. It allows to combine queries on a number of structure-related databases not covered by other projects at present. Thus, information on both many and few protein structures can be used efficiently. The web interface for COLUMBA is available at http://www.columba-db.de.

Base Sequence↗

What makes a gene name? Named entity recognition in the biomedical literature.

The recognition of biomedical concepts in natural text (named entity recognition, NER) is a key technology for automatic or semi-automatic analysis of textual resources. Precise NER tools are a prerequisite for many applications working on text, such as information retrieval, information extraction or document classification. Over the past years, the problem has achieved considerable attention in the bioinformatics community and experience has shown that NER in the life sciences is a rather difficult problem. Several systems and algorithms have been devised and implemented. In this paper, the problems and resources in NER research are described, the principal algorithms underlying most systems sketched, and the current state-of-the-art in the field surveyed.

Algorithms↗

Finding kinetic parameters using text mining.

The mathematical modeling and description of complex biological processes has become more and more important over the last years. Systems biology aims at the computational simulation of complex systems, up to whole cell simulations. An essential part focuses on solving a large number of parameterized differential equations. However, measuring those parameters is an expensive task, and finding them in the literature is very laborious. We developed a text mining system that supports researchers in their search for experimentally obtained parameters for kinetic models. Our system classifies full text documents regarding the question whether or not they contain appropriate data using a support vector machine. We evaluated our approach on a manually tagged corpus of 800 documents and found that it outperforms keyword searches in abstracts by a factor of five in terms of precision.

Computational Biology↗