PubMed Health⌕ Search

Biomedical subjects

Alistair G Rust

Publications and source records attributed to Alistair G Rust.

9 recordsLinked to original sources

Systems biology approaches identify ATF3 as a negative regulator of Toll-like receptor 4.

The innate immune system is absolutely required for host defence, but, uncontrolled, it leads to inflammatory disease. This control is mediated, in part, by cytokines that are secreted by macrophages. Immune regulation is extraordinarily complex, and can be best investigated with systems approaches (that is, using computational tools to predict regulatory networks arising from global, high-throughput data sets). Here we use cluster analysis of a comprehensive set of transcriptomic data derived from Toll-like receptor (TLR)-activated macrophages to identify a prominent group of genes that appear to be regulated by activating transcription factor 3 (ATF3), a member of the CREB/ATF family of transcription factors. Network analysis predicted that ATF3 is part of a transcriptional complex that also contains members of the nuclear factor (NF)-kappaB family of transcription factors. Promoter analysis of the putative ATF3-regulated gene cluster demonstrated an over-representation of closely apposed ATF3 and NF-kappaB binding sites, which was verified by chromatin immunoprecipitation and hybridization to a DNA microarray. This cluster included important cytokines such as interleukin (IL)-6 and IL-12b. ATF3 and Rel (a component of NF-kappaB) were shown to bind to the regulatory regions of these genes upon macrophage activation. A kinetic model of Il6 and Il12b messenger RNA expression as a function of ATF3 and NF-kappaB promoter binding predicted that ATF3 is a negative regulator of Il6 and Il12b transcription, and this hypothesis was validated using Atf3-null mice. ATF3 seems to inhibit Il6 and Il12b transcription by altering chromatin structure, thereby restricting access to transcription factors. Because ATF3 is itself induced by lipopolysaccharide, it seems to regulate TLR-stimulated inflammatory responses as part of a negative-feedback loop.

Activating Transcription Factor 3↗

Transcription binding site prediction using Markov models.

One of the main goals of analysing DNA sequences is to understand the temporal and positional information that specifies gene expression. An important step in this process is the recognition of gene expression regulatory elements. Experimental procedures for this are slow and costly. In this paper we present a computational non-supervised algorithm that facilitates the process by statistically identifying the most likely regions within a putative regulatory sequence. A probabilistic technique is presented, based on the approximation of regulatory DNA with a Markov chain, for the location of putative transcription factor binding sites in a single stretch of DNA. Hereto we developed a procedure to approximate the order of Markov model for a given DNA sequence that circumvents some of the prohibitive assumptions underlying Markov modeling. Application of the algorithm to data from 55 genes in five species shows the high sensitivity of this Markov search algorithm. Our algorithm does not require any prior knowledge in the form of description or cross-genomic comparison; it is context sensitive and takes DNA heterogeneity into account.

Artificial Intelligence↗

Improving computational predictions of cis-regulatory binding sites.

The location of cis-regulatory binding sites determine the connectivity of genetic regulatory networks and therefore constitute a natural focal point for research into the many biological systems controlled by such regulatory networks. Accurate computational prediction of these binding sites would facilitate research into a multitude of key areas, including embryonic development, evolution, pharmacogenemics, cancer and many other transcriptional diseases, and is likely to be an important precursor for the reverse engineering of genome wide, genetic regulatory networks. Many algorithmic strategies have been developed for the computational prediction of cis-regulatory binding sites but currently all approaches are prone to high rates of false positive predictions, and many are highly dependent on additional information, limiting their usefulness as research tools. In this paper we present an approach for improving the accuracy of a selection of established prediction algorithms. Firstly, it is shown that species specific optimization of algorithmic parameters can, in some cases, significantly improve the accuracy of algorithmic predictions. Secondly, it is demonstrated that the use of non-linear classification algorithms to integrate predictions from multiple sources can result in more accurate predictions. Finally, it is shown that further improvements in prediction accuracy can be gained with the use of biologically inspired post-processing of predictions.

Algorithms↗

A data integration methodology for systems biology.

Different experimental technologies measure different aspects of a system and to differing depth and breadth. High-throughput assays have inherently high false-positive and false-negative rates. Moreover, each technology includes systematic biases of a different nature. These differences make network reconstruction from multiple data sets difficult and error-prone. Additionally, because of the rapid rate of progress in biotechnology, there is usually no curated exemplar data set from which one might estimate data integration parameters. To address these concerns, we have developed data integration methods that can handle multiple data sets differing in statistical power, type, size, and network coverage without requiring a curated training data set. Our methodology is general in purpose and may be applied to integrate data from any existing and future technologies. Here we outline our methods and then demonstrate their performance by applying them to simulated data sets. The results show that these methods select true-positive data elements much more accurately than classical approaches. In an accompanying companion paper, we demonstrate the applicability of our approach to biological data. We have integrated our methodology into a free open source software package named POINTILLIST.

Informatics↗

A data integration methodology for systems biology: experimental verification.

The integration of data from multiple global assays is essential to understanding dynamic spatiotemporal interactions within cells. In a companion paper, we reported a data integration methodology, designated Pointillist, that can handle multiple data types from technologies with different noise characteristics. Here we demonstrate its application to the integration of 18 data sets relating to galactose utilization in yeast. These data include global changes in mRNA and protein abundance, genome-wide protein-DNA interaction data, database information, and computational predictions of protein-DNA and protein-protein interactions. We divided the integration task to determine three network components: key system elements (genes and proteins), protein-protein interactions, and protein-DNA interactions. Results indicate that the reconstructed network efficiently focuses on and recapitulates the known biology of galactose utilization. It also provided new insights, some of which were verified experimentally. The methodology described here, addresses a critical need across all domains of molecular and cell biology, to effectively integrate large and disparate data sets.

Chromatin Immunoprecipitation↗

New computational approaches for analysis of cis-regulatory networks.

The investigation and modeling of gene regulatory networks requires computational tools specific to the task. We present several locally developed software tools that have been used in support of our ongoing research into the embryogenesis of the sea urchin. These tools are especially well suited to iterative refinement of models through experimental and computational investigation. They include: BioArray, a macroarray spot processing program; SUGAR, a system to display and correlate large-BAC sequence analyses; SeqComp and FamilyRelations, programs for comparative sequence analysis; and NetBuilder, an environment for creating and analyzing models of gene networks. We also present an overview of the process used to build our model of the Strongylocentrotus purpuratus endomesoderm gene network. Several of the tools discussed in this paper are still in active development and some are available as open source.

Chromosomes, Artificial, Bacterial↗

A provisional regulatory gene network for specification of endomesoderm in the sea urchin embryo.

We present the current form of a provisional DNA sequence-based regulatory gene network that explains in outline how endomesodermal specification in the sea urchin embryo is controlled. The model of the network is in a continuous process of revision and growth as new genes are added and new experimental results become available; see http://www.its.caltech.edu/~mirsky/endomeso.htm (End-mes Gene Network Update) for the latest version. The network contains over 40 genes at present, many newly uncovered in the course of this work, and most encoding DNA-binding transcriptional regulatory factors. The architecture of the network was approached initially by construction of a logic model that integrated the extensive experimental evidence now available on endomesoderm specification. The internal linkages between genes in the network have been determined functionally, by measurement of the effects of regulatory perturbations on the expression of all relevant genes in the network. Five kinds of perturbation have been applied: (1) use of morpholino antisense oligonucleotides targeted to many of the key regulatory genes in the network; (2) transformation of other regulatory factors into dominant repressors by construction of Engrailed repressor domain fusions; (3) ectopic expression of given regulatory factors, from genetic expression constructs and from injected mRNAs; (4) blockade of the beta-catenin/Tcf pathway by introduction of mRNA encoding the intracellular domain of cadherin; and (5) blockade of the Notch signaling pathway by introduction of mRNA encoding the extracellular domain of the Notch receptor. The network model predicts the cis-regulatory inputs that link each gene into the network. Therefore, its architecture is testable by cis-regulatory analysis. Strongylocentrotus purpuratus and Lytechinus variegatus genomic BAC recombinants that include a large number of the genes in the network have been sequenced and annotated. Tests of the cis-regulatory predictions of the model are greatly facilitated by interspecific computational sequence comparison, which affords a rapid identification of likely cis-regulatory elements in advance of experimental analysis. The network specifies genomically encoded regulatory processes between early cleavage and gastrula stages. These control the specification of the micromere lineage and of the initial veg(2) endomesodermal domain; the blastula-stage separation of the central veg(2) mesodermal domain (i.e., the secondary mesenchyme progenitor field) from the peripheral veg(2) endodermal domain; the stabilization of specification state within these domains; and activation of some downstream differentiation genes. Each of the temporal-spatial phases of specification is represented in a subelement of the network model, that treats regulatory events within the relevant embryonic nuclei at particular stages.

Animals↗

Genome annotation techniques: new approaches and challenges.

As more of the human genome draft sequence is finished, and genomes from other organisms begin to be sequenced, the demand for accurate and reliable genome annotation will increase significantly. To facilitate this industrial-scale genome annotation, automated bioinformatics solutions are increasingly required. As a result, automatic genome annotation systems have become more important in gene discovery within recent years. The design of such large-scale bioinformatics systems is an evolving and dynamic field, based on central cores of bioinformatics software tools and relational databases. Not only must these systems efficiently manage and integrate large volumes of genomic data, but they must also deliver accurate gene predictions and effectively distribute annotation data to the biosciences community.

Computational Biology↗

A genomic regulatory network for development.

Development of the body plan is controlled by large networks of regulatory genes. A gene regulatory network that controls the specification of endoderm and mesoderm in the sea urchin embryo is summarized here. The network was derived from large-scale perturbation analyses, in combination with computational methodologies, genomic data, cis-regulatory analysis, and molecular embryology. The network contains over 40 genes at present, and each node can be directly verified at the DNA sequence level by cis-regulatory analysis. Its architecture reveals specific and general aspects of development, such as how given cells generate their ordained fates in the embryo and why the process moves inexorably forward in developmental time.

Animals↗