PubMed Health⌕ Search

Biomedical subjects

Joachim Selbig

Publications and source records attributed to Joachim Selbig.

At least 19 recordsLinked to original sources

Visualization and analysis of molecular data.

This chapter provides an overview of visualization and analysis techniques applied to large-scale datasets from genomics, metabolomics, and proteomics. The aim is to reduce the number of variables (genes, metabolites, or proteins) by extracting a small set of new relevant variables, usually termed components. The advantages and disadvantages of the classical principal component analysis (PC A) are discussed and a link is given to the closely related singular value decomposition and multidimensional scaling. Special emphasis is given to the recent trend toward the use of independent component analysis, which aims to extract statistically independent components and, therefore, provides usually more meaningful components than PCA. We also discuss normalization techniques and their influence on the result of different analytical techniques.

Algorithms↗

A gentle guide to the analysis of metabolomic data.

Modern molecular biology crucially relies on computational tools to handle and interpret the large amounts of data that are generated by high-throughput measurements. To this end, much effort is dedicated to devise novel sophisticated methods that allow one to integrate, evaluate, and analyze biological data. However, prior to an application of specifically designed methods, simple and well-known statistical approaches often provide a more appropriate starting point for further analysis. This chapter seeks to describe several well-established approaches to data analysis, including various clustering techniques, discriminant function analysis, principal component analysis, multidimensional scaling, and classification trees. The chapter is accompanied by a webpage, describing the application of all algorithms in a ready-to-use format.

Arabidopsis↗

Validation and functional annotation of expression-based clusters based on gene ontology.

BACKGROUND: The biological interpretation of large-scale gene expression data is one of the paramount challenges in current bioinformatics. In particular, placing the results in the context of other available functional genomics data, such as existing bio-ontologies, has already provided substantial improvement for detecting and categorizing genes of interest. One common approach is to look for functional annotations that are significantly enriched within a group or cluster of genes, as compared to a reference group. RESULTS: In this work, we suggest the information-theoretic concept of mutual information to investigate the relationship between groups of genes, as given by data-driven clustering, and their respective functional categories. Drawing upon related approaches (Gibbons and Roth, Genome Research 12:1574-1581, 2002), we seek to quantify to what extent individual attributes are sufficient to characterize a given group or cluster of genes. CONCLUSION: We show that the mutual information provides a systematic framework to assess the relationship between groups or clusters of genes and their functional annotations in a quantitative way. Within this framework, the mutual information allows us to address and incorporate several important issues, such as the interdependence of functional annotations and combinatorial combinations of attributes. It thus supplements and extends the conventional search for overrepresented attributes within a group or cluster of genes. In particular taking combinations of attributes into account, the mutual information opens the way to uncover specific functional descriptions of a group of genes or clustering result. All datasets and functional annotations used in this study are publicly available. All scripts used in the analysis are provided as additional files.

Algorithms↗

Structural kinetic modeling of metabolic networks.

To develop and investigate detailed mathematical models of metabolic processes is one of the primary challenges in systems biology. However, despite considerable advance in the topological analysis of metabolic networks, kinetic modeling is still often severely hampered by inadequate knowledge of the enzyme-kinetic rate laws and their associated parameter values. Here we propose a method that aims to give a quantitative account of the dynamical capabilities of a metabolic system, without requiring any explicit information about the functional form of the rate equations. Our approach is based on constructing a local linear model at each point in parameter space, such that each element of the model is either directly experimentally accessible or amenable to a straightforward biochemical interpretation. This ensemble of local linear models, encompassing all possible explicit kinetic models, then allows for a statistical exploration of the comprehensive parameter space. The method is exemplified on two paradigmatic metabolic systems: the glycolytic pathway of yeast and a realistic-scale representation of the photosynthetic Calvin cycle.

Computer Simulation↗

Bioinformatics approach to predicting HIV drug resistance.

The emergence of drug resistance remains one of the most challenging issues in the treatment of HIV-1 infection. The extreme replication dynamics of HIV facilitates its escape from the selective pressure exerted by the human immune system and by the applied combination drug therapy. This article reviews computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genotypic and phenotypic data. Genotypic assays are based on the analysis of mutations associated with reduced drug susceptibility, but are difficult to interpret due to the numerous mutations and mutational patterns that confer drug resistance. Phenotypic resistance or susceptibility can be experimentally evaluated by measuring the inhibition of the viral replication in cell culture assays. However, this procedure is expensive and time consuming.

Anti-HIV Agents↗

Evolution of HIV resistance during treatment interruption in experienced patients and after restarting a new therapy.

BACKGROUND: To analyse the evolution of resistance patterns in patients undergoing treatment interruption (TI) and re-initiating highly active anti-retroviral therapy (HAART). METHODS: HIV-RT and -PR gene-sequences were analysed in 14 patients (>5 failing prior drugs) before and during TI and under a new HAART. Genotypes were interpreted using two bioinformatics systems. Additionally, virus load (VL) and CD4(+)-T-cell counts were measured. RESULTS: Six patients (42%) achieved sustained undetectable VL up to one year after TI (responders), while 8 (57%) maintained VL of more than 2,000 copies/mL (non-responders). Different patterns of resistance-mutations evolution were detected. During TI loss of all mutations was observed in three patients, a reduction of mutations was detected in seven patients, and no alteration was seen in four patients. In the responders, 87.5% of protease inhibitor (PI)-resistance mutations waned during TI and remained undetectable under the new treatment. In contrast, in the non-responder group most PI-resistance mutations continued noticeable under the new therapy. Loss of primary PI-resistance mutations and the presence of one fully active PI in the new regimen significantly correlated with success of subsequent treatment (p=0.028). In two patients new reverse transcriptase associated mutations were detected during TI, G190A (NNRTI mutation) and K70R (NRTI mutation). Appearance of K70R could be explained by a reverse direction of a previously described pathway of thymidin analogues mutation resistance development, while G190A could be due to prolonged subinhibitory drug levels after cessation of NNRTIs. CONCLUSION: In the evolution of HAART-resistance, different patterns were observed in responders and non-responders during but not before TI. Absence of PI-resistance associated mutations during and after TI and administration of a predicted fully active PI for the new therapy correlated with success. Newly detected mutations during TI may indicate reversibility of previously described mutational pathways.

Anti-Retroviral Agents↗

Computational methods for the design of effective therapies against drug resistant HIV strains.

The development of drug resistance is a major obstacle to successful treatment of HIV infection. The extraordinary replication dynamics of HIV facilitates its escape from selective pressure exerted by the human immune system and by combination drug therapy. We have developed several computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genomic data.

Database Management Systems↗

Non-linear PCA: a missing data approach.

MOTIVATION: Visualizing and analysing the potential non-linear structure of a dataset is becoming an important task in molecular biology. This is even more challenging when the data have missing values. RESULTS: Here, we propose an inverse model that performs non-linear principal component analysis (NLPCA) from incomplete datasets. Missing values are ignored while optimizing the model, but can be estimated afterwards. Results are shown for both artificial and experimental datasets. In contrast to linear methods, non-linear methods were able to give better missing value estimations for non-linear structured data. APPLICATION: We applied this technique to a time course of metabolite data from a cold stress experiment on the model plant Arabidopsis thaliana, and could approximate the mapping function from any time point to the metabolite responses. Thus, the inverse NLPCA provides greatly improved information for better understanding the complex response to cold stress. CONTACT: scholz@mpimp-golm.mpg.de.

Adaptation, Physiological↗

Species-specific analysis of protein sequence motifs using mutual information.

BACKGROUND: Protein sequence motifs are by definition short fragments of conserved amino acids, often associated with a specific function. Accordingly protein sequence profiles derived from multiple sequence alignments provide an alternative description of functional motifs characterizing families of related sequences. Such profiles conveniently reflect functional necessities by pointing out proximity at conserved sequence positions as well as depicting distances at variable positions. Discovering significant conservation characteristics within the variable positions of profiles mirrors group-specific and, in particular, evolutionary features of the underlying sequences. RESULTS: We describe the tool PROfile analysis based on Mutual Information (PROMI) that enables comparative analysis of user-classified protein sequences. PROMI is implemented as a web service using Perl and R as well as other publicly available packages and tools on the server-side. On the client-side platform-independence is achieved by generally applied internet delivery standards. As one possible application analysis of the zinc finger C2H2-type protein domain is introduced to illustrate the functionality of the tool. CONCLUSION: The web service PROMI should assist researchers to detect evolutionary correlations in protein profiles of defined biological sequences. It is available at http://promi.mpimp-golm.mpg.de where additional documentation can be found.

Amino Acid Motifs↗

Estimating HIV evolutionary pathways and the genetic barrier to drug resistance.

BACKGROUND: The evolution of drug-resistant viruses challenges the management of human immunodeficiency virus (HIV) infections. Understanding this evolutionary process is important for the design of effective therapeutic strategies. METHODS: We used mutagenetic trees, a family of probabilistic graphical models, to describe the accumulation of resistance-associated mutations in the viral genome. On the basis of these models, we defined the genetic barrier, a quantity that summarizes the difficulty for the virus to escape from the selective pressure of the drug by developing escape mutations. RESULTS: From HIV reverse-transcriptase sequences that had been obtained from treated patients, we derived evolutionary models for zidovudine, zidovudine plus lamivudine, and zidovudine plus didanosine. The genetic barriers to resistance to zidovudine, stavudine, lamivudine, and didanosine, for the above 3 regimens, were computed and analyzed. We found both the mode and the rate of development of resistance to be heterogeneous. The genetic barrier to zidovudine resistance was increased if lamivudine was added to zidovudine but was decreased for didanosine. The barrier to lamivudine resistance was maintained with zidovudine plus didanosine, whereas the barrier to didanosine resistance was reduced most with zidovudine plus lamivudine. CONCLUSION: Mutagenetic trees provide a quantitative picture of the evolution of drug resistance. The genetic barrier is a useful tool for design of effective treatment strategies.

Anti-HIV Agents↗

Mtreemix: a software package for learning and using mixture models of mutagenetic trees.

SUMMARY: Mixture models of mutagenetic trees constitute a class of probabilistic models for describing evolutionary processes that are characterized by the accumulation of permanent genetic changes. They have been applied to model the accumulation of chromosomal gains and losses in tumor development and the development of drug resistance-associated mutations in the HIV genome.Mtreemix is a software package for estimating mutagenetic trees mixture models from observed cross-sectional data and for using these models for predictions. We provide programs for model fitting, model selection, simulation, likelihood computation and waiting time estimation. AVAILABILITY: Mtreemix, including source code, documentation, sample data files and precompiled Solaris and Linux binaries, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de/

Algorithms↗

Extension of the visualization tool MapMan to allow statistical analysis of arrays, display of corresponding genes, and comparison with known responses.

MapMan is a user-driven tool that displays large genomics datasets onto diagrams of metabolic pathways or other processes. Here, we present new developments, including improvements of the gene assignments and the user interface, a strategy to visualize multilayered datasets, the incorporation of statistics packages, and extensions of the software to incorporate more biological information including visualization of corresponding genes and horizontal searches for similar global responses across large numbers of arrays.

Genes, Plant↗

A Robot-based platform to measure multiple enzyme activities in Arabidopsis using a set of cycling assays: comparison of changes of enzyme activities and transcript levels during diurnal cycles and in prolonged darkness.

A platform has been developed to measure the activity of 23 enzymes that are involved in central carbon and nitrogen metabolism in Arabidopsis thaliana. Activities are assayed in optimized stopped assays and the product then determined using a suite of enzyme cycling assays. The platform requires inexpensive equipment, is organized in a modular manner to optimize logistics, calculates results automatically, combines high sensitivity with throughput, can be robotized, and has a throughput of three to four activities in 100 samples per person/day. Several of the assays, including those for sucrose phosphate synthase, ADP glucose pyrophosphorylase (AGPase), ferredoxin-dependent glutamate synthase, glycerokinase, and shikimate dehydrogenase, provide large advantages over previous approaches. This platform was used to analyze the diurnal changes of enzyme activities in wild-type Columbia-0 (Col-0) and the starchless plastid phosphoglucomutase (pgm) mutant, and in Col-0 during a prolongation of the night. The changes of enzyme activities were compared with the changes of transcript levels determined with the Affymetrix ATH1 array. Changes of transcript levels typically led to strongly damped changes of enzyme activity. There was no relation between the amplitudes of the diurnal changes of transcript and enzyme activity. The largest diurnal changes in activity were found for AGPase and nitrate reductase. Examination of the data and comparison with the literature indicated that these are mainly because of posttranslational regulation. The changes of enzyme activity are also strongly delayed, with the delay varying from enzyme to enzyme. It is proposed that enzyme activities provide a quasi-stable integration of regulation at several levels and provide useful data for the characterization and diagnosis of different physiological states. As an illustration, a decision tree constructed using data from Col-0 during diurnal changes and a prolonged dark treatment was used to show that, irrespective of the time of harvest during the diurnal cycle, the pgm mutant resembles a wild-type plant that has been exposed to a 3 d prolongation of the night.

Arabidopsis↗

Estimating mutual information using B-spline functions--an improved similarity measure for analysing gene expression data.

BACKGROUND: The information theoretic concept of mutual information provides a general framework to evaluate dependencies between variables. In the context of the clustering of genes with similar patterns of expression it has been suggested as a general quantity of similarity to extend commonly used linear measures. Since mutual information is defined in terms of discrete variables, its application to continuous data requires the use of binning procedures, which can lead to significant numerical errors for datasets of small or moderate size. RESULTS: In this work, we propose a method for the numerical estimation of mutual information from continuous data. We investigate the characteristic properties arising from the application of our algorithm and show that our approach outperforms commonly used algorithms: The significance, as a measure of the power of distinction from random correlation, is significantly increased. This concept is subsequently illustrated on two large-scale gene expression datasets and the results are compared to those obtained using other similarity measures.A C++ source code of our algorithm is available for non-commercial use from kloska@scienion.de upon request. CONCLUSION: The utilisation of mutual information as similarity measure enables the detection of non-linear correlations in gene expression datasets. Frequently applied linear correlation measures, which are often used on an ad-hoc basis without further justification, are thereby extended.

Algorithms↗

PaVESy: Pathway Visualization and Editing System.

UNLABELLED: A data managing system for editing and visualization of biological pathways is presented. The main component of PaVESy (Pathway Visualization and Editing System) is a relational SQL database system. The database design allows storage of biological objects, such as metabolites, proteins, genes and respective relations, which are required to assemble metabolic and regulatory biological interactions. The database model accommodates highly flexible annotation of biological objects by user-defined attributes. In addition, specific roles of objects are derived from these attributes in the context of user-defined interactions, e.g. in the course of pathway generation or during editing of the database content. Furthermore, the user may organize and arrange the database content within a folder structure and is free to group and annotate database objects of interest within customizable subsets. Thus, we allow an individualized view on the database content and facilitate user customization. A JAVA-based class library was developed, which serves as the database programming interface to PaVESy. This API provides classes, which implement the concepts of object persistence in SQL databases, such as entries, interactions, annotations, folders and subsets. We created editing and visualization tools for navigation in and visualization of the database content. User approved pathway assemblies are stored and may be retrieved for continued modification, annotation and export. Data export is interfaced with a range of network visualization programs, such as Pajek or other software allowing import of SBML or GML data format. AVAILABILITY: http://pavsey.mpimp-golm.mpg.de

Database Management Systems↗

Hypothesis-driven approach to predict transcriptional units from gene expression data.

MOTIVATION: A major issue in computational biology is the reconstruction of functional relationships among genes, for example the definition of regulatory or biochemical pathways. One step towards this aim is the elucidation of transcriptional units, which are characterized by co-responding changes in mRNA expression levels. These units of genes will allow the generation of hypotheses about respective functional interrelationships. Thus, the focus of analysis currently moves from well-established functional assignment through comparison of protein and DNA sequences towards analysis of transcriptional co-response. Tools that allow deducing common control of gene expression have the potential to complement and extend routine BLAST comparisons, because gene function may be inferred from common transcriptional control. RESULTS: We present a co-clustering strategy of genome sequence information and gene expression data, which was applied to identify transcriptional units within diverse compendia of expression profiles. The phenomenon of prokaryotic operons was selected as an ideal test case to generate well-founded hypotheses about transcriptional units. The existence of overlapping and ambiguous operon definitions allowed the investigation of constitutive and conditional expression of transcriptional units in independent gene expression experiments of Escherichia coli. Our approach allowed identification of operons with high accuracy. Furthermore, both constitutive mRNA co-response as well as conditional differences became apparent. Thus, we were able to generate insight into the possible biological relevance of gene co-response. We conclude that the suggested strategy will be amenable in general to the identification of transcriptional units beyond the chosen example of E.coli operons. AVAILABILITY: The analyses of E.coli transcript data presented here are available upon request or at http://csbdb.mpimp-golm.mpg.de/

Algorithms↗

MAPMAN: a user-driven tool to display genomics data sets onto diagrams of metabolic pathways and other biological processes.

MAPMAN is a user-driven tool that displays large data sets onto diagrams of metabolic pathways or other processes. SCAVENGER modules assign the measured parameters to hierarchical categories (formed 'BINs', 'subBINs'). A first build of TRANSCRIPTSCAVENGER groups genes on the Arabidopsis Affymetrix 22K array into >200 hierarchical categories, providing a breakdown of central metabolism (for several pathways, down to the single enzyme level), and an overview of secondary metabolism and cellular processes. METABOLITESCAVENGER groups hundreds of metabolites into pathways or groups of structurally related compounds. An IMAGEANNOTATOR module uses these groupings to organise and display experimental data sets onto diagrams of the users' choice. A modular structure allows users to edit existing categories, add new categories and develop SCAVENGER modules for other sorts of data. MAPMAN is used to analyse two sets of 22K Affymetrix arrays that investigate the response of Arabidopsis rosettes to low sugar: one investigates the response to a 6-h extension of the night, and the other compares wild-type Columbia-0 (Col-0) and the starchless pgm mutant (plastid phosphoglucomutase) at the end of the night. There were qualitatively similar responses in both treatments. Many genes involved in photosynthesis, nutrient acquisition, amino acid, nucleotide, lipid and cell wall synthesis, cell wall modification, and RNA and protein synthesis were repressed. Many genes assigned to amino acid, nucleotide, lipid and cell wall breakdown were induced. Changed expression of genes for trehalose metabolism point to a role for trehalose-6-phosphate (Tre6P) as a starvation signal. Widespread changes in the expression of genes encoding receptor kinases, transcription factors, components of signalling pathways, proteins involved in post-translational modification and turnover, and proteins involved in the synthesis and sensing of cytokinins, abscisic acid (ABA) and ethylene revealing large-scale rewiring of the regulatory network is an early response to sugar depletion.

Abscisic Acid↗

MetaGeneAlyse: analysis of integrated transcriptional and metabolite data.

UNLABELLED: New techniques in sample preparation allow high throughput analysis of samples on the transcriptional as well as on the metabolic level. We present a service accessible via the web that allows the analysis of integrated data sets that combine gene-expression data and metabolic data. After uploading, data sets can be normalized, clustered by various methods and results can be graphically visualized. All calculations are carried out on a server, so even time- and memory-consuming analyses can be done independently of the performance of the client. AVAILABILITY: The service is accessible via web-interface at http://metagenealyse.mpimp-golm.mpg.de/

Algorithms↗