PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “regulatory module”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Sequence turnover and tandem repeats in cis-regulatory modules in drosophila.

The path by which regulatory sequence can change, yet preserve function, is an important open question for both evolution and bioinformatics. The recent sequencing of two additional species of Drosophila plus the wealth of data on gene regulation in the fruit fly provides new means for addressing this question. For regulatory sequences, indels account for more base pairs (bp) of change than substitutions (between Drosophila melanogaster and Drosophila yakuba), though they are fewer in number. Using Drosophila pseudoobscura as an out-group, we can distinguish insertions from deletions (with maximum parsimony criteria), and find a ratio between 1 and 5 (insertions to deletions) that is species dependent and much larger than the ratio of 1/8 for neutral sequences (Petrov and Hartl 1998). Because neutral sequence is rapidly cleared from the genome, most noncoding regions which preserve their length between D. melanogaster-D. pseudoobscura and have an excess of insertions over deletions should be functional. A fraction of 15%-18% (i.e., more than 20 standard deviations from random expectation) of the regulatory sequence is covered by low copy number tandem repeats whose repeating unit has an average length of 5-10 bp and which occur preferentially (25%-45% coverage) in indels. All indels may be due to tandem repeats if we extrapolate the detection efficiency of the repeat-finding algorithms using the observed point mutation rate between the species we compare. Sequence creation by local duplication accords with the tendency for multiple copies of transcription factor-binding sites to occur in regulatory modules. Thus, indel events and tandem repeats in particular need to be incorporated into models of regulatory evolution because they can alter the rate at which beneficial variants arise and should also influence bioinformatic algorithms that parse regulatory sequences into binding sites.

Animals↗

Finding regulatory modules through large-scale gene-expression data analysis.

MOTIVATION: The use of gene microchips has enabled a rapid accumulation of gene-expression data. One of the major challenges of analyzing this data is the diversity, in both size and signal strength, of the various modules in the gene regulatory networks of organisms. RESULTS: Based on the iterative signature algorithm [Bergmann,S., Ihmels,J. and Barkai,N. (2002) Phys. Rev. E 67, 031902], we present an algorithm-the progressive iterative signature algorithm (PISA)-that, by sequentially eliminating modules, allows unsupervised identification of both large and small regulatory modules. We applied PISA to a large set of yeast gene-expression data, and, using the Gene Ontology database as a reference, found that the algorithm is much better able to identify regulatory modules than methods based on high-throughput transcription-factor binding experiments or on comparative genomics.

Algorithms↗

CREME: a framework for identifying cis-regulatory modules in human-mouse conserved segments.

MOTIVATION: The binding of transcription factors to specific regulatory sequence elements is a primary mechanism for controlling gene transcription. Recent findings suggest a modular organization of binding sites for transcription factors that cooperate in the regulation of genes. In this work we establish a framework for finding recurrent cis-regulatory modules in the promoters of a selected set of genes and scoring their statistical significance. RESULTS: Proceeding from a database of identified binding site motifs and their genomic locations we seek motifs whose frequency in the selected promoters is different than in a background promoter set. We present several statistical tests designed for this purpose. We provide a hashing algorithm for detecting combinations of these motifs that co-occur in clusters within the selected promoters. The significance of such co-occurrences is evaluated using novel statistical scores. Our methods are combined in CREME, a suite of software which includes a browser for viewing the pattern of occurrence of selected cis-regulatory modules. We applied our methodology to find modules within human-mouse conserved promoter segments, focusing on cell cycle regulated genes and stress response related genes. To validate the biological significance of the identified modules we tested whether the associated genes tended to be co-expressed or share similar function. In the cell cycle set five of the seven identified sets of genes were coherently expressed. On the stress response data four of the six detected sets fell predominantly into well-defined functional sub-categories.

Algorithms↗

Computational detection of genomic cis-regulatory modules applied to body patterning in the early Drosophila embryo.

BACKGROUND: Regulation of gene transcription is crucial for the function and development of all organisms. While gene prediction programs that identify protein coding sequence are used with remarkable success in the annotation of genomes, the development of computational methods to analyze noncoding regions and to delineate transcriptional control elements is still in its infancy. RESULTS: Here we present novel algorithms to detect cis-regulatory modules through genome wide scans for clusters of transcription factor binding sites using three levels of prior information. When binding sites for the factors are known, our statistical segmentation algorithm, Ahab, yields about 150 putative gap gene regulated modules, with no adjustable parameters other than a window size. If one or more related modules are known, but no binding sites, repeated motifs can be found by a customized Gibbs sampler and input to Ahab, to predict genes with similar regulation. Finally using only the genome, we developed a third algorithm, Argos, that counts and scores clusters of overrepresented motifs in a window of sequence. Argos recovers many of the known modules, upstream of the segmentation genes, with no training data. CONCLUSIONS: We have demonstrated, in the case of body patterning in the Drosophila embryo, that our algorithms allow the genome-wide identification of regulatory modules. We believe that Ahab overcomes many problems of recent approaches and we estimated the false positive rate to be about 50%. Argos is the first successful attempt to predict regulatory modules using only the genome without training data. Complete results and module predictions across the Drosophila genome are available at http://uqbar.rockefeller.edu/~siggia/.

Algorithms↗

Identification of novel regulatory modules in dicotyledonous plants using expression data and comparative genomics.

BACKGROUND: Transcriptional regulation plays an important role in the control of many biological processes. Transcription factor binding sites (TFBSs) are the functional elements that determine transcriptional activity and are organized into separable cis-regulatory modules, each defining the cooperation of several transcription factors required for a specific spatio-temporal expression pattern. Consequently, the discovery of novel TFBSs in promoter sequences is an important step to improve our understanding of gene regulation. RESULTS: Here, we applied a detection strategy that combines features of classic motif overrepresentation approaches in co-regulated genes with general comparative footprinting principles for the identification of biologically relevant regulatory elements and modules in Arabidopsis thaliana, a model system for plant biology. In total, we identified 80 TFBSs and 139 regulatory modules, most of which are novel, and primarily consist of two or three regulatory elements that could be linked to different important biological processes, such as protein biosynthesis, cell cycle control, photosynthesis and embryonic development. Moreover, studying the physical properties of some specific regulatory modules revealed that Arabidopsis promoters have a compact nature, with cooperative TFBSs located in close proximity of each other. CONCLUSION: These results create a starting point to unravel regulatory networks in plants and to study the regulation of biological processes from a systems biology point of view.

Arabidopsis↗

Cell-specific regulatory modules control expression of genes in vascular and visceral smooth muscle tissues.

A novel approach with chimeric SM22alpha/telokin promoters was used to identify gene regulatory modules that are required for regulating the expression of genes in distinct smooth muscle tissues. Conventional deletion or mutation analysis of promoters does not readily distinguish regulatory elements that are required for basal gene expression from those required for expression in specific smooth muscle tissues. In the present study, the mouse telokin gene was isolated, and a 370-bp (-190 to 180) minimal promoter was identified that directs visceral smooth muscle-specific expression in vivo in transgenic mice. The visceral smooth muscle-specific expression of the telokin promoter transgene is in marked contrast to the reported arterial smooth muscle-specific expression of a 536-bp minimal SM22alpha (-475 to 61) promoter transgene. To begin to identify regulatory elements that are responsible for the distinct tissue-specific expression of these promoters, a chimeric promoter in which a 172-bp SM22alpha gene fragment (-288 to -116) was fused to the minimal telokin promoter was generated and characterized. The -288 to -116 SM22alpha gene fragment significantly increased telokin promoter activity in vascular smooth muscle cells in vitro and in vivo. Conversely, a fragment of the telokin promoter (-94 to -49) increased the activity of the SM22alpha promoter in visceral smooth muscle cells of the bladder. Together, these data demonstrate that both vascular- and visceral smooth muscle-specific regulatory modules direct gene expression in subsets of smooth muscle tissues.

AT Rich Sequence↗

Re-arranging the Cis-regulatory Modules of Hox Complex in Drosophila via FLP-FRT and CRISPR/Cas9.

FLP-FRT, a well-established technique for genome manipulation, and the revolutionary CRISPR/Cas9, known for its targeted indels, are combined in a novel approach. This unique method is applied to the Hox genes in the Drosophila melanogaster bithorax complex, which are closely located to the cis-regulatory modules that define their spatial-temporal regulation. The number and position of these genes are directly correlated to their expression pattern. This chapter unveils the exciting potential of this combinatorial use of FLP-FRT and CRISPR-Cas9 to rearrange the cis-regulatory modules of the Hox complex in Drosophila melanogaster.

Animals↗

Regulatory modules shared within gene classes as well as across gene classes can be detected by the same in silico approach.

Transcriptional regulation depends on the binding of transcription factors to their corresponding binding sites. The response to cellular signals is often mediated by the cooperative binding of transcription factors to well defined regulatory modules consisting of at least two transcription factor binding sites. Such regulatory modules can be responsible for the common regulation of genes within a gene class or confer a common function to promoters belonging to different gene classes. We developed in silico models representing a common framework of potential regulatory sites specific for one promoter class (actins). We also generated models for two different functional promoter modules both of which confer responsiveness to tumor necrosis factor (TNF) and interferon (IFN) to a variety of promoters. All models exhibited high selectivity, e.g. the mammalian muscle actin promoter model produced no false negatives in a database search.

Actins↗

Computational reconstruction of transcriptional regulatory modules of the yeast cell cycle.

BACKGROUND: A transcriptional regulatory module (TRM) is a set of genes that is regulated by a common set of transcription factors (TFs). By organizing the genome into TRMs, a living cell can coordinate the activities of many genes and carry out complex functions. Therefore, identifying TRMs is helpful for understanding gene regulation. RESULTS: Integrating gene expression and ChIP-chip data, we develop a method, called MOdule Finding Algorithm (MOFA), for reconstructing TRMs of the yeast cell cycle. MOFA identified 87 TRMs, which together contain 336 distinct genes regulated by 40 TFs. Using various kinds of data, we validated the biological relevance of the identified TRMs. Our analysis shows that different combinations of a fairly small number of TFs are responsible for regulating a large number of genes involved in different cell cycle phases and that there may exist crosstalk between the cell cycle and other cellular processes. MOFA is capable of finding many novel TF-target gene relationships and can determine whether a TF is an activator or/and a repressor. Finally, MOFA refines some clusters proposed by previous studies and provides a better understanding of how the complex expression program of the cell cycle is regulated. CONCLUSION: MOFA was developed to reconstruct TRMs of the yeast cell cycle. Many of these TRMs are in agreement with previous studies. Further, MOFA inferred many interesting modules and novel TF combinations. We believe that computational analysis of multiple types of data will be a powerful approach to studying complex biological systems when more and more genomic resources such as genome-wide protein activity data and protein-protein interaction data become available.

Algorithms↗

Stubb: a program for discovery and analysis of cis-regulatory modules.

Given the DNA-binding specificities (motifs) of one or more transcription factors, an important bioinformatics problem is to discover significant clusters of binding sites for the transcription factors(s). Such clusters often correspond to cis-regulatory modules mediating regulation of an adjacent gene. In earlier work, we developed the Stubb program that uses a probabilistic model and a maximum likelihood approach to efficiently detect cis-regulatory modules over genomic scales. It may optionally exploit a second related genome to improve module prediction accuracy. We describe here the use of a web-based interface for the Stubb program. The interface is equipped with a special post-processing step for in-depth analysis of specific modules, in order to reveal individual binding sites predicted in the module. The web server may be accessed at the URL http://stubb.rockefeller.edu/.

Algorithms↗

Cross-species comparison significantly improves genome-wide prediction of cis-regulatory modules in Drosophila.

BACKGROUND: The discovery of cis-regulatory modules in metazoan genomes is crucial for understanding the connection between genes and organism diversity. It is important to quantify how comparative genomics can improve computational detection of such modules. RESULTS: We run the Stubb software on the entire D. melanogaster genome, to obtain predictions of modules involved in segmentation of the embryo. Stubb uses a probabilistic model to score sequences for clustering of transcription factor binding sites, and can exploit multiple species data within the same probabilistic framework. The predictions are evaluated using publicly available gene expression data for thousands of genes, after careful manual annotation. We demonstrate that the use of a second genome (D. pseudoobscura) for cross-species comparison significantly improves the prediction accuracy of Stubb, and is a more sensitive approach than intersecting the results of separate runs over the two genomes. The entire list of predictions is made available online. CONCLUSION: Evolutionary conservation of modules serves as a filter to improve their detection in silico. The future availability of additional fruitfly genomes therefore carries the prospect of highly specific genome-wide predictions using Stubb.

Algorithms↗

De novo cis-regulatory module elicitation for eukaryotic genomes.

Transcription regulation is controlled by coordinated binding of one or more transcription factors in the promoter regions of genes. In many species, especially higher eukaryotes, transcription factor binding sites tend to occur as homotypic or heterotypic clusters, also known as cis-regulatory modules. The number of sites and distances between the sites, however, vary greatly in a module. We propose a statistical model to describe the underlying cluster structure as well as individual motif conservation and develop a Monte Carlo motif screening strategy for predicting novel regulatory modules in upstream sequences of coregulated genes. We demonstrate the power of the method with examples ranging from bacterial to insect and human genomes.

Base Sequence↗

Genome-wide computational prediction of transcriptional regulatory modules reveals new insights into human gene expression.

The identification of regulatory regions is one of the most important and challenging problems toward the functional annotation of the human genome. In higher eukaryotes, transcription-factor (TF) binding sites are often organized in clusters called cis-regulatory modules (CRM). While the prediction of individual TF-binding sites is a notoriously difficult problem, CRM prediction has proven to be somewhat more reliable. Starting from a set of predicted binding sites for more than 200 TF families documented in Transfac, we describe an algorithm relying on the principle that CRMs generally contain several phylogenetically conserved binding sites for a few different TFs. The method allows the prediction of more than 118,000 CRMs within the human genome. A subset of these is shown to be bound in vivo by TFs using ChIP-chip. Their analysis reveals, among other things, that CRM density varies widely across the genome, with CRM-rich regions often being located near genes encoding transcription factors involved in development. Predicted CRMs show a surprising enrichment near the 3' end of genes and in regions far from genes. We document the tendency for certain TFs to bind modules located in specific regions with respect to their target genes and identify TFs likely to be involved in tissue-specific regulation. The set of predicted CRMs, which is made available as a public database called PReMod (http://genomequebec.mcgill.ca/PReMod), will help analyze regulatory mechanisms in specific biological systems.

Algorithms↗

A computational genomics approach to identify cis-regulatory modules from chromatin immunoprecipitation microarray data--a case study using E2F1.

Advances in high-throughput technologies, such as ChIP-chip, and the completion of human and mouse genomic sequences now allow analysis of the mechanisms of gene regulation on a systems level. In this study, we have developed a computational genomics approach (termed ChIPModules), which begins with experimentally determined binding sites and integrates positional weight matrices constructed from transcription factor binding sites, a comparative genomics approach, and statistical learning methods to identify transcriptional regulatory modules. We began with E2F1 binding site information obtained from ChIP-chip analyses of ENCODE regions, from both HeLa and MCF7 cells. Our approach not only distinguished targets from nontargets with a high specificity, but it also identified five regulatory modules for E2F1. One of the identified modules predicted a colocalization of E2F1 and AP-2alpha on a set of target promoters with an intersite distance of <270 bp. We tested this prediction using ChIP-chip assays with arrays containing approximately 14,000 human promoters. We found that both E2F1 and AP-2alpha bind within the predicted distance to a large number of human promoters, demonstrating the strength of our sequence-based, unbiased, and universal protocol. Finally, we have used our ChIPModules approach to develop a database that includes thousands of computationally identified and/or experimentally verified E2F1 target promoters.

Base Pairing↗

Drosophila mef2 expression during mesoderm development is controlled by a complex array of cis-acting regulatory modules.

The function of the Drosophila mef2 gene, a member of the MADS box supergene family of transcription factors, is critical for terminal differentiation of the three major muscle cell types, namely somatic, visceral, and cardiac. During embryogenesis, mef2 undergoes multiple phases of expression, which are characterized by initial broad mesodermal expression, followed by restricted expression in the dorsal mesoderm, specific expression in muscle progenitors, and sustained expression in the differentiated musculatures. In this study, evidence is presented that temporally and spatially specific mef2 expression is controlled by a complex array of cis-acting regulatory modules that are responsive to different genetic signals. Functional testing of approximately 12 kb of 5' flanking region of the mef2 gene showed that the initial widespread mesodermal expression is achieved through a 280-bp twist-dependent enhancer. The subsequent dorsal mesoderm-restricted mef2 expression is mediated through a 460-bp dpp-responsive regulatory module, which involves the function of the Smad4 homolog Medea and contains several binding sites for Medea and Mad. The analysis also showed that regulated mef2 expression in the caudal and trunk visceral mesoderm, which give rise to longitudinal and circular gut musculatures, respectively, is under the control of distinct enhancer elements. In addition, mef2 expression in the cardioblasts of the heart is dependent upon at least two distinct enhancers, which are active at different periods during embryogenesis. Moreover, multiple regulatory elements are differentially activated for specific expression in presumptive muscle founders, prefusion myoblasts, and differentiated muscle fibers. Taken together, the presented data suggest that specific expression of the mef2 gene in myogenic lineages in the Drosophila embryo is the result of multiple genetic inputs that act in an additive manner upon distinct enhancers in the 5' flanking region.

Animals↗

CisView: a browser and database of cis-regulatory modules predicted in the mouse genome.

To facilitate the analysis of gene regulatory regions of the mouse genome, we developed a CisView (http://lgsun.grc.nia.nih.gov/cisview), a browser and database of genome-wide potential transcription factor binding sites (TFBSs) that were identified using 134 position-weight matrices and 219 sequence patterns from various sources and were presented with the information about sequence conservation, neighboring genes and their structures, GO annotations, protein domains, DNA repeats and CpG islands. Analysis of the distribution of TFBSs revealed that many TFBSs (N = 145) were over-represented near transcription start sites. We also identified potential cis-regulatory modules (CRMs) defined as clusters of conserved TFBSs in the entire mouse genome. Out of 739 074 CRMs, 157 442 had a significantly higher regulatory potential score than semi-random sequences generated with a 3rd-order Markov process. The CisView browser provides a user-friendly computer environment for studying transcription regulation on a whole-genome scale and can also be used for interpreting microarray experiments and identifying putative targets of transcription factors.

Animals↗

Segmenting the fly embryo: a logical analysis of the pair-rule cross-regulatory module.

This manuscript reports a dynamical analysis of the pair-rule cross-regulatory module controlling segmentation in Drosophila melanogaster. We propose a logical model accounting for the ability of the pair-rule module to determine the formation of alternate juxtaposed Engrailed- and Wingless-expressing cells that form the (para)segmental boundaries. This module has the intrinsic capacity to generate four distinct expression states, each characterized by the expression of a particular combination of pair-rule genes or expression mode. The selection of one of these expression modes depends on the maternal and gap inputs, but also crucially on cross-regulations among pair-rule genes. The latter are instrumental in the interpretation of the maternal-gap pre-pattern. Our logical model allows the qualitative reproduction of the patterns of pair-rule gene expressions corresponding to the wild type situation, to loss-of-function and cis-regulatory mutations, and to ectopic pair-rule expressions. Furthermore, this model provides a formal explanation for the morphogenetic role of the initial bell-shaped expression of the gene even-skipped, i.e. for the distinct effects of different levels of the Even-skipped protein on its target pair-rule genes. It also accounts for the requirement of Even-skipped for the formation of all Engrailed-stripes. Finally, it provides new insights into the roles and evolutionary origins of the apparent redundancies in the regulatory architecture of the pair-rule module.

Animals↗

Identifying cis-regulatory modules by combining comparative and compositional analysis of DNA.

MOTIVATION: Predicting cis-regulatory modules (CRMs) in higher eukaryotes is a challenging computational task. Commonly used methods to predict CRMs based on the signal of transcription factor binding sites (TFBS) are limited by prior information about transcription factor specificity. More general methods that bypass the reliance on TFBS models are needed for comprehensive CRM prediction. RESULTS: We have developed a method to predict CRMs called CisPlusFinder that identifies high density regions of perfect local ungapped sequences (PLUSs) based on multiple species conservation. By assuming that PLUSs contain core TFBS motifs that are locally overrepresented, the method attempts to capture the expected features of CRM structure and evolution. Applied to a benchmark dataset of CRMs involved in early Drosophila development, CisPlusFinder predicts more annotated CRMs than all other methods tested. Using the REDfly database, we find that some 'false positive' predictions in the benchmark dataset correspond to recently annotated CRMs. Our work demonstrates that CRM prediction methods that combine comparative genomic data with statistical properties of DNA may achieve reasonable performance when applied genome-wide in the absence of an a priori set of known TFBS motifs. AVAILABILITY: The program CisPlusFinder can be downloaded at http://jakob.genetik.uni-koeln.de/bioinformatik/people/nora/nora.html. All software is licensed under the Lesser GNU Public License (LGPL).

Algorithms↗