PubMed Health⌕ Search

Biomedical subjects

Rainer Breitling

Publications and source records attributed to Rainer Breitling.

At least 19 recordsLinked to original sources

Synteny plot quality control with SyntenyQC.

SUMMARY: SyntenyQC is a data pre-processing tool for the construction of synteny plots. It supports genomic data collection, annotation and dereplication to facilitate (and in some cases fundamentally enable) the construction of informative synteny plots. AVAILABILITY AND IMPLEMENTATION: SyntenyQC is a command line app developed using Python version 3.10 and tested using pytest. SyntenyQC is available on PyPI (https://pypi.org/project/SyntenyQC) under the MIT License, along with a detailed user tutorial. Package tests can be viewed at https://github.com/Tim-Kirkwood/SyntenyQC.

Synteny↗

Precision mapping of the metabolome.

The global study of the structure and dynamics of metabolic networks has been hindered by a lack of techniques that identify metabolites and their biochemical relationship in complex mixtures. The recent application of Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) to metabolomic analysis suggests a way to tackle the problem. A lower-cost alternative to high-field FTICR-MS, the Orbitrap mass analyzer, promises accelerated activity in this area. Here, we show how the ultra-high mass accuracy and resolution provided by this new generation of mass spectrometers can help to identify metabolites and connect them into metabolic networks. Data from perturbation studies and isotope-tracking experiments can complement this information to create metabolic maps de novo and chart unexplored areas of metabolism.

Algorithms↗

RankProd: a bioconductor package for detecting differentially expressed genes in meta-analysis.

UNLABELLED: While meta-analysis provides a powerful tool for analyzing microarray experiments by combining data from multiple studies, it presents unique computational challenges. The Bioconductor package RankProd provides a new and intuitive tool for this purpose in detecting differentially expressed genes under two experimental conditions. The package modifies and extends the rank product method proposed by Breitling et al., [(2004) FEBS Lett., 573, 83-92] to integrate multiple microarray studies from different laboratories and/or platforms. It offers several advantages over t-test based methods and accepts pre-processed expression datasets produced from a wide variety of platforms. The significance of the detection is assessed by a non-parametric permutation test, and the associated P-value and false discovery rate (FDR) are included in the output alongside the genes that are detected by user-defined criteria. A visualization plot is provided to view actual expression levels for each gene with estimated significance measurements. AVAILABILITY: RankProd is available at Bioconductor http://www.bioconductor.org. A web-based interface will soon be available at http://cactus.salk.edu/RankProd

Computational Biology↗

Biological microarray interpretation: the rules of engagement.

Gene expression microarrays are now established as a standard tool in biological and biochemical laboratories. Interpreting the masses of data generated by this technology poses a number of unusual new challenges. Over the past few years a consensus has begun to emerge concerning the most important pitfalls and the proper ways to avoid them. This review provides an overview of these ideas, beginning with relevant aspects of experimental design and normalization, but focusing in particular on the various tools and concepts that help to interpret microarray results. These new approaches make it much easier to extract biologically relevant and reliable hypotheses in an objective and reasonably unbiased fashion.

Animals↗

A lock-and-key model for protein-protein interactions.

MOTIVATION: Protein-protein interaction networks are one of the major post-genomic data sources available to molecular biologists. They provide a comprehensive view of the global interaction structure of an organism's proteome, as well as detailed information on specific interactions. Here we suggest a physical model of protein interactions that can be used to extract additional information at an intermediate level: It enables us to identify proteins which share biological interaction motifs, and also to identify potentially missing or spurious interactions. RESULTS: Our new graph model explains observed interactions between proteins by an underlying interaction of complementary binding domains (lock-and-key model). This leads to a novel graph-theoretical algorithm to identify bipartite subgraphs within protein-protein interaction networks where the underlying data are taken from yeast two-hybrid experimental results. By testing on synthetic data, we demonstrate that under certain modelling assumptions, the algorithm will return correct domain information about each protein in the network. Tests on data from various model organisms show that the local and global patterns predicted by the model are indeed found in experimental data. Using functional and protein structure annotations, we show that bipartite subnetworks can be identified that correspond to biologically relevant interaction motifs. Some of these are novel and we discuss an example involving SH3 domains from the Saccharomyces cerevisiae interactome. AVAILABILITY: The algorithm (in Matlab format) is available (see http://www.maths.strath.ac.uk/~aas96106/lock_key.html).

Algorithms↗

Regulation of ubiquitin-binding proteins by monoubiquitination.

Proteins containing ubiquitin-binding domains (UBDs) interact with ubiquitinated targets and regulate diverse biological processes, including endocytosis, signal transduction, transcription and DNA repair. Many of the UBD-containing proteins are also themselves monoubiquitinated, but the functional role and the mechanisms that underlie this modification are less well understood. Here, we demonstrate that monoubiquitination of the endocytic proteins Sts1, Sts2, Eps15 and Hrs results in intramolecular interactions between ubiquitin and their UBDs, thereby preventing them from binding in trans to ubiquitinated targets. Permanent monoubiquitination of these proteins, mimicked by the fusion of ubiquitin to their carboxyl termini, impairs their ability to regulate trafficking of ubiquitinated receptors. Moreover, we mapped the in vivo monoubiquitination site in Sts2 and demonstrated that its mutation enhances the Sts2-mediated effects of epidermal-growth-factor-receptor downregulation. We propose that monoubiquitination of ubiquitin-binding proteins inhibits their capacity to bind to and control the functions of ubiquitinated targets in vivo.

Adaptor Proteins, Signal Transducing↗

Network theory to understand microarray studies of complex diseases.

Complex diseases, such as allergy, diabetes and obesity depend on altered interactions between multiple genes, rather than changes in a single causal gene. DNA microarray studies of a complex disease often implicate hundreds of genes in the pathogenesis. This indicates that many different mechanisms and pathways are involved. How can we understand such complexity? How can hypotheses be formulated and tested? One approach is to organize the data in network models and to analyze these in a top-down manner. Globally, networks in nature are often characterized by a small number of highly connected nodes, while the majority of nodes have few connections. The highly connected nodes serve as hubs that affect many other nodes. Such hubs have key roles in the network. In yeast cells, for example, deletion of highly connected proteins is associated with increased lethality, compared to deletion of less connected proteins. This suggests the biological relevance of networks. Moving down in the network structure, there may be sub-networks or modules with specific functions. These modules may be further dissected to analyze individual nodes. In the context of DNA microarray studies of complex diseases, gene-interaction networks may contain modules of co-regulated or interacting genes that have distinct biological functions. Such modules may be linked to specific gene polymorphisms, transcription factors, cellular functions and disease mechanisms. Genes that are reliably active only in the context of their modules can be considered markers for the activity of the modules and may thus be promising candidates for biomarkers or therapeutic targets. This review aims to give an introduction to network theory and how it can be applied to microarray studies of complex diseases.

Humans↗

Current challenges in quantitative modeling of epidermal growth factor signaling.

Over the last decade, epidermal growth factor (EGF) signaling has been used repeatedly as a test-bed for pioneering computational systems biology. Recent breakthroughs in our molecular understanding of EGF signaling pose new challenges for mathematical modeling strategies. Three key areas emerge as particularly relevant: the pervasive importance of compartmentalization and endosomal trafficking; the complexity of signalosome complexes; and the regulatory influence of diffusion and spatiality. Each one of them demands a drastic change in current computational approaches. We discuss recent developments in the field that address these emerging aspects in a new generation of more realistic - and potential more useful - models of EGF signaling.

Animals↗

GeneRank: using search engine technology for the analysis of microarray experiments.

BACKGROUND: Interpretation of simple microarray experiments is usually based on the fold-change of gene expression between a reference and a "treated" sample where the treatment can be of many types from drug exposure to genetic variation. Interpretation of the results usually combines lists of differentially expressed genes with previous knowledge about their biological function. Here we evaluate a method--based on the PageRank algorithm employed by the popular search engine Google--that tries to automate some of this procedure to generate prioritized gene lists by exploiting biological background information. RESULTS: GeneRank is an intuitive modification of PageRank that maintains many of its mathematical properties. It combines gene expression information with a network structure derived from gene annotations (gene ontologies) or expression profile correlations. Using both simulated and real data we find that the algorithm offers an improved ranking of genes compared to pure expression change rankings. CONCLUSION: Our modification of the PageRank algorithm provides an alternative method of evaluating microarray experimental results which combines prior knowledge about the underlying network. GeneRank offers an improvement compared to assessing the importance of a gene based on its experimentally observed fold-change alone and may be used as a basis for further analytical developments.

Algorithms↗

Vector analysis as a fast and easy method to compare gene expression responses between different experimental backgrounds.

BACKGROUND: Gene expression studies increasingly compare expression responses between different experimental backgrounds (genetic, physiological, or phylogenetic). By focusing on dynamic responses rather than a direct comparison of static expression levels, this type of study allows a finer dissection of primary and secondary regulatory effects in the various backgrounds. Usually, results of such experiments are presented in the form of Venn diagrams, which are intuitive and visually appealing, but lack a statistical foundation. RESULTS: Here we introduce Vector Analysis (VA) as a simple, yet principled, approach to comparing expression responses in different experimental backgrounds. VA enables the automatic assignment of genes to response prototypes and provides statistical significance estimates to eliminate spurious response patterns. The application of VA to a real dataset, comparing nutrient starvation responses in wild type and mutant Arabidopsis plants, reveals that consistent patterns of expression behavior are present in the data and are reliably detected by the algorithm. CONCLUSION: Vector analysis is a flexible, easy-to-use technique to compare gene expression patterns in different experimental backgrounds. It compares favorably with the classical Venn diagram approach and can be implemented manually using spreadsheets, such as Excel, or automatically by using the supplied software.

Algorithms↗

Biological master games: using biologists' reasoning to guide algorithm development for integrated functional genomics.

We review some powerful new algorithms that build on the intuitive biological interpretation techniques for statistical analysis of functional genomics experiments. Although they were originally designed for transcriptomics, we argue that these algorithms are applicable to any type of -omics study (transcriptomics, proteomics, metabolomics). Rank Products (RP), a strictly non-parametric test statistic to detect differentially regulated elements (genes, proteins, metabolites) in genome-wide screens. RP is particularly powerful for noisy data and low numbers of replicates and makes full use of the availability of a large number of parallel measurements that is typical of modern large-scale experiments. Iterative Group Analysis (iGA), a statistical method that makes the transition from regulated single elements to significant classes of elements, and thus provides an automatic functional annotation of an experiment. Graph-based iGA (GiGA), an extension of iGA that combines experimental data with a broad variety of biological annotations to highlight physiologically relevant regions in a given "evidence graph" (e.g., metabolic networks, signaling pathway diagrams, protein interaction maps). The sequential application of these techniques yields an increasingly abstract interpretation of experimental data that is at the same time quantitative, statistically rigorous, and biologically significant. The results can be used either as helpful tools to guide data visualization and exploration, or as the input for downstream computational applications in a systems biology framework.

Algorithms↗

FrankSum: new feature selection method for protein function prediction.

In the study of in silico functional genomics, improving the performance of protein function prediction is the ultimate goal for identifying proteins associated with defined cellular functions. The classical prediction approach is to employ pairwise sequence alignments. However this method often faces difficulties when no statistically significant homologous sequences are identified. An alternative way is to predict protein function from sequence-derived features using machine learning. In this case the choice of possible features which can be derived from the sequence is of vital importance to ensure adequate discrimination to predict function. In this paper we have successfully selected biologically significant features for protein function prediction. This was performed using a new feature selection method (FrankSum) that avoids data distribution assumptions, uses a data independent measurement (p-value) within the feature, identifies redundancy between features and uses an appropriate ranking criterion for feature selection. We have shown that classifiers generated from features selected by FrankSum outperforms classifiers generated from full feature sets, randomly selected features and features selected from the Wrapper method. We have also shown the features are concordant across all species and top ranking features are biologically informative. We conclude that feature selection is vital for successful protein function prediction and FrankSum is one of the feature selection methods that can be applied successfully to such a domain.

Amino Acid Sequence↗

Rank-based methods as a non-parametric alternative of the T-statistic for the analysis of biological microarray data.

We have recently introduced a rank-based test statistic, RankProducts (RP), as a new non-parametric method for detecting differentially expressed genes in microarray experiments. It has been shown to generate surprisingly good results with biological datasets. The basis for this performance and the limits of the method are, however, little understood. Here we explore the performance of such rank-based approaches under a variety of conditions using simulated microarray data, and compare it with classical Wilcoxon rank sums and t-statistics, which form the basis of most alternative differential gene expression detection techniques. We show that for realistic simulated microarray datasets, RP is more powerful and accurate for sorting genes by differential expression than t-statistics or Wilcoxon rank sums - in particular for replicate numbers below 10, which are most commonly used in biological experiments. Its relative performance is particularly strong when the data are contaminated by non-normal random noise or when the samples are very inhomogenous, e.g. because they come from different time points or contain a mixture of affected and unaffected cells. However, RP assumes equal measurement variance for all genes and tends to give overly optimistic p-values when this assumption is violated. It is therefore essential that proper variance stabilizing normalization is performed on the data before calculating the RP values. Where this is impossible, another rank-based variant of RP (average ranks) provides a useful alternative with very similar overall performance. The Perl scripts implementing the simulation and evaluation are available upon request. Implementations of the RP method are available for download from the authors website (http://www.brc.dcs.gla.ac.uk/glama).

Algorithms↗

Feature selection and the class imbalance problem in predicting protein function from sequence.

When the standard approach to predict protein function by sequence homology fails, other alternative methods can be used that require only the amino acid sequence for predicting function. One such approach uses machine learning to predict protein function directly from amino acid sequence features. However, there are two issues to consider before successful functional prediction can take place: identifying discriminatory features, and overcoming the challenge of a large imbalance in the training data. We show that by applying feature subset selection followed by undersampling of the majority class, significantly better support vector machine (SVM) classifiers are generated compared with standard machine learning approaches. As well as revealing that the features selected could have the potential to advance our understanding of the relationship between sequence and function, we also show that undersampling to produce fully balanced data significantly improves performance. The best discriminating ability is achieved using SVMs together with feature selection and full undersampling; this approach strongly outperforms other competitive learning algorithms. We conclude that this combined approach can generate powerful machine learning classifiers for predicting protein function directly from sequence.

Algorithms↗

The potassium-dependent transcriptome of Arabidopsis reveals a prominent role of jasmonic acid in nutrient signaling.

Full genome microarrays were used to assess transcriptional responses of Arabidopsis seedlings to changing external supply of the essential macronutrient potassium (K(+)). Rank product statistics and iterative group analysis were employed to identify differentially regulated genes and statistically significant coregulated sets of functionally related genes. The most prominent response was found for genes linked to the phytohormone jasmonic acid (JA). Transcript levels for the JA biosynthetic enzymes lipoxygenase, allene oxide synthase, and allene oxide cyclase were strongly increased during K(+) starvation and quickly decreased after K(+) resupply. A large number of well-known JA responsive genes showed the same expression profile, including genes involved in storage of amino acids (VSP), glucosinolate production (CYP79), polyamine biosynthesis (ADC2), and defense (PDF1.2). Our findings highlight a novel role of JA in nutrient signaling and stress management through a variety of physiological processes such as nutrient storage, recycling, and reallocation. Other highly significant K(+)-responsive genes discovered in our study encoded cell wall proteins (e.g. extensins and arabinogalactans) and ion transporters (e.g. the high-affinity K(+) transporter HAK5 and the nitrate transporter NRT2.1) as well as proteins with a putative role in Ca(2+) signaling (e.g. calmodulins). On the basis of our results, we propose candidate genes involved in K(+) perception and signaling as well as a network of molecular processes underlying plant adaptation to K(+) deficiency.

Arabidopsis↗

Rank products: a simple, yet powerful, new method to detect differentially regulated genes in replicated microarray experiments.

One of the main objectives in the analysis of microarray experiments is the identification of genes that are differentially expressed under two experimental conditions. This task is complicated by the noisiness of the data and the large number of genes that are examined simultaneously. Here, we present a novel technique for identifying differentially expressed genes that does not originate from a sophisticated statistical model but rather from an analysis of biological reasoning. The new technique, which is based on calculating rank products (RP) from replicate experiments, is fast and simple. At the same time, it provides a straightforward and statistically stringent way to determine the significance level for each gene and allows for the flexible control of the false-detection rate and familywise error rate in the multiple testing situation of a microarray experiment. We use the RP technique on three biological data sets and show that in each case it performs more reliably and consistently than the non-parametric t-test variant implemented in Tusher et al.'s significance analysis of microarrays (SAM). We also show that the RP results are reliable in highly noisy data. An analysis of the physiological function of the identified genes indicates that the RP approach is powerful for identifying biologically relevant expression changes. In addition, using RP can lead to a sharp reduction in the number of replicate experiments needed to obtain reproducible results.

Acute Disease↗

Graph-based iterative Group Analysis enhances microarray interpretation.

BACKGROUND: One of the most time-consuming tasks after performing a gene expression experiment is the biological interpretation of the results by identifying physiologically important associations between the differentially expressed genes. A large part of the relevant functional evidence can be represented in the form of graphs, e.g. metabolic and signaling pathways, protein interaction maps, shared GeneOntology annotations, or literature co-citation relations. Such graphs are easily constructed from available genome annotation data. The problem of biological interpretation can then be described as identifying the subgraphs showing the most significant patterns of gene expression. We applied a graph-based extension of our iterative Group Analysis (iGA) approach to obtain a statistically rigorous identification of the subgraphs of interest in any evidence graph. RESULTS: We validated the Graph-based iterative Group Analysis (GiGA) by applying it to the classic yeast diauxic shift experiment of DeRisi et al., using GeneOntology and metabolic network information. GiGA reliably identified and summarized all the biological processes discussed in the original publication. Visualization of the detected subgraphs allowed the convenient exploration of the results. The method also identified several processes that were not presented in the original paper but are of obvious relevance to the yeast starvation response. CONCLUSIONS: GiGA provides a fast and flexible delimitation of the most interesting areas in a microarray experiment, and leads to a considerable speed-up and improvement of the interpretation process.

Algorithms↗

Biologically valid linear factor models of gene expression.

MOTIVATION: The identification of physiological processes underlying and generating the expression pattern observed in microarray experiments is a major challenge. Principal component analysis (PCA) is a linear multivariate statistical method that is regularly employed for that purpose as it provides a reduced-dimensional representation for subsequent study of possible biological processes responding to the particular experimental conditions. Making explicit the data assumptions underlying PCA highlights their lack of biological validity thus making biological interpretation of the principal components problematic. A microarray data representation which enables clear biological interpretation is a desirable analysis tool. RESULTS: We address this issue by employing the probabilistic interpretation of PCA and proposing alternative linear factor models which are based on refined biological assumptions. A practical study on two well-understood microarray datasets highlights the weakness of PCA and the greater biological interpretability of the linear models we have developed.

Algorithms↗