PubMed Health⌕ Search

Biomedical subjects

Harmen J Bussemaker

Publications and source records attributed to Harmen J Bussemaker.

18 recordsLinked to original sources

Hotspots of transcription factor colocalization in the genome of Drosophila melanogaster.

Regulation of gene expression is a highly complex process that requires the concerted action of many proteins, including sequence-specific transcription factors, cofactors, and chromatin proteins. In higher eukaryotes, the interplay between these proteins and their interactions with the genome still is poorly understood. We systematically mapped the in vivo binding sites of seven transcription factors with diverse physiological functions, five cofactors, and two heterochromatin proteins at approximately 1-kb resolution in a 2.9 Mb region of the Drosophila melanogaster genome. Surprisingly, all tested transcription factors and cofactors show strongly overlapping localization patterns, and the genome contains many "hotspots" that are targeted by all of these proteins. Several control experiments show that the strong overlap is not an artifact of the techniques used. Colocalization hotspots are 1-5 kb in size, spaced on average by approximately 50 kb, and preferentially located in regions of active transcription. We provide evidence that protein-protein interactions play a role in the hotspot association of some transcription factors. Colocalization hotspots constitute a previously uncharacterized type of feature in the genome of Drosophila, and our results provide insights into the general targeting mechanisms of transcription regulators in a higher eukaryote.

Animals↗

Statistical mechanical modeling of genome-wide transcription factor occupancy data by MatrixREDUCE.

MOTIVATION: Regulation of gene expression by a transcription factor requires physical interaction between the factor and the DNA, which can be described by a statistical mechanical model. Based on this model, we developed the MatrixREDUCE algorithm, which uses genome-wide occupancy data for a transcription factor (e.g. ChIP-chip) and associated nucleotide sequences to discover the sequence-specific binding affinity of the transcription factor. Advantages of our approach are that the information for all probes on the microarray is efficiently utilized because there is no need to delineate "bound" and "unbound" sequences, and that, unlike information content-based methods, it does not require a background sequence model. RESULTS: We validated the performance of MatrixREDUCE by inferring the sequence-specific binding affinities for several transcription factors in S. cerevisiae and comparing the results with three other independent sources of transcription factor sequence-specific affinity information: (i) experimental measurement of transcription factor binding affinities for specific oligonucleotides, (ii) reporter gene assays for promoters with systematically mutated binding sites, and (iii) relative binding affinities obtained by modeling transcription factor-DNA interactions based on co-crystal structures of transcription factors bound to DNA substrates. We show that transcription factor binding affinities inferred by MatrixREDUCE are in good agreement with all three validating methods. AVAILABILITY: MatrixREDUCE source code is freely available for non-commercial use at http://www.bussemakerlab.org/. The software runs on Linux, Unix, and Mac OS X.

Algorithms↗

Detecting transcriptionally active regions using genomic tiling arrays.

We have developed a method for interpreting genomic tiling array data, implemented as the program TranscriptionDetector. Probed loci expressed above background are identified by combining replicates in a way that makes minimal assumptions about the data. We performed medium-resolution Anopheles gambiae tiling array experiments and found extensive transcription of both coding and non-coding regions. Our method also showed improved detection of transcriptional units when applied to high-density tiling array data for ten human chromosomes.

Animals↗

Profiling condition-specific, genome-wide regulation of mRNA stability in yeast.

The steady-state abundance of an mRNA is determined by the balance between transcription and decay. Although regulation of transcription has been well studied both experimentally and computationally, regulation of transcript stability has received little attention. We developed an algorithm, MatrixREDUCE, that discovers the position-specific affinity matrices for unknown RNA-binding factors and infers their condition-specific activities, using only genomic sequence data and steady-state mRNA expression data as input. We identified and computationally characterized the binding sites for six mRNA stability regulators in Saccharomyces cerevisiae, which include two members of the Pumilio-homology domain (Puf) family of RNA-binding proteins, Puf3p and Puf4p. We provide computational and experimental evidence that regulation of mRNA stability by these factors is modulated in response to a variety of environmental stimuli.

Gene Expression Profiling↗

T-profiler: scoring the activity of predefined groups of genes using gene expression data.

One of the key challenges in the analysis of gene expression data is how to relate the expression level of individual genes to the underlying transcriptional programs and cellular state. Here we describe T-profiler, a tool that uses the t-test to score changes in the average activity of predefined groups of genes. The gene groups are defined based on Gene Ontology categorization, ChIP-chip experiments, upstream matches to a consensus transcription factor binding motif or location on the same chromosome. If desired, an iterative procedure can be used to select a single, optimal representative from sets of overlapping gene groups. T-profiler makes it possible to interpret microarray data in a way that is both intuitive and statistically rigorous, without the need to combine experiments or choose parameters. Currently, gene expression data from Saccharomyces cerevisiae and Candida albicans are supported. Users can upload their microarray data for analysis on the web at http://www.t-profiler.org.

Algorithms↗

Comparative genome sequencing of Drosophila pseudoobscura: chromosomal, gene, and cis-element evolution.

We have sequenced the genome of a second Drosophila species, Drosophila pseudoobscura, and compared this to the genome sequence of Drosophila melanogaster, a primary model organism. Throughout evolution the vast majority of Drosophila genes have remained on the same chromosome arm, but within each arm gene order has been extensively reshuffled, leading to a minimum of 921 syntenic blocks shared between the species. A repetitive sequence is found in the D. pseudoobscura genome at many junctions between adjacent syntenic blocks. Analysis of this novel repetitive element family suggests that recombination between offset elements may have given rise to many paracentric inversions, thereby contributing to the shuffling of gene order in the D. pseudoobscura lineage. Based on sequence similarity and synteny, 10,516 putative orthologs have been identified as a core gene set conserved over 25-55 million years (Myr) since the pseudoobscura/melanogaster divergence. Genes expressed in the testes had higher amino acid sequence divergence than the genome-wide average, consistent with the rapid evolution of sex-specific proteins. Cis-regulatory sequences are more conserved than random and nearby sequences between the species--but the difference is slight, suggesting that the evolution of cis-regulatory elements is flexible. Overall, a pattern of repeat-mediated chromosomal rearrangement, and high coadaptation of both male genes and cis-regulatory sequences emerges as important themes of genome divergence between these species of Drosophila.

Animals↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

Defining transcriptional networks through integrative modeling of mRNA expression and transcription factor binding data.

BACKGROUND: Functional genomics studies are yielding information about regulatory processes in the cell at an unprecedented scale. In the yeast S. cerevisiae, DNA microarrays have not only been used to measure the mRNA abundance for all genes under a variety of conditions but also to determine the occupancy of all promoter regions by a large number of transcription factors. The challenge is to extract useful information about the global regulatory network from these data. RESULTS: We present MA-Networker, an algorithm that combines microarray data for mRNA expression and transcription factor occupancy to define the regulatory network of the cell. Multivariate regression analysis is used to infer the activity of each transcription factor, and the correlation across different conditions between this activity and the mRNA expression of a gene is interpreted as regulatory coupling strength. Applying our method to S. cerevisiae, we find that, on average, 58% of the genes whose promoter region is bound by a transcription factor are true regulatory targets. These results are validated by an analysis of enrichment for functional annotation, response for transcription factor deletion, and over-representation of cis-regulatory motifs. We are able to assign directionality to transcription factors that control divergently transcribed genes sharing the same promoter region. Finally, we identify an intrinsic limitation of transcription factor deletion experiments related to the combinatorial nature of transcriptional control, to which our approach provides an alternative. CONCLUSION: Our reliable classification of ChIP positives into functional and non-functional TF targets based on their expression pattern across a wide range of conditions provides a starting point for identifying the unknown sequence features in non-coding DNA that directly or indirectly determine the context dependence of transcription factor action. Complete analysis results are available for browsing or download at http://bussemaker.bio.columbia.edu/papers/MA-Networker/.

Algorithms↗

Distinct HP1 and Su(var)3-9 complexes bind to sets of developmentally coexpressed genes depending on chromosomal location.

Heterochromatin proteins are thought to play key roles in chromatin structure and gene regulation, yet very few genes have been identified that are regulated by these proteins. We performed large-scale mapping and analysis of in vivo target loci of the proteins HP1, HP1c, and Su(var)3-9 in Drosophila Kc cells, which are of embryonic origin. For each protein, we identified approximately 100-200 target genes among >6000 probed loci. We found that HP1 and Su(var)3-9 bind together to transposable elements and genes that are predominantly pericentric. In addition, Su(var)3-9 binds without HP1 to a distinct set of nonpericentric genes. On chromosome 4, HP1 binds to many genes, mostly independent of Su(var)3-9. The binding pattern of HP1c is largely different from those of HP1 and Su(var)3-9. Target genes of HP1 and Su(var)3-9 show lower expression levels in Kc cells than do nontarget genes, but not if they are located in pericentric regions. Strikingly, in pericentric regions, target genes of Su(var)3-9 and HP1 are predominantly embryo-specific genes, whereas on the chromosome arms Su(var)3-9 is preferentially associated with a set of male-specific genes. These results demonstrate that, depending on chromosomal location, the HP1 and Su(var)3-9 proteins form different complexes that associate with specific sets of developmentally coexpressed genes.

Animals↗

The human transcriptome map reveals extremes in gene density, intron length, GC content, and repeat pattern for domains of highly and weakly expressed genes.

The chromosomal gene expression profiles established by the Human Transcriptome Map (HTM) revealed a clustering of highly expressed genes in about 30 domains, called ridges. To physically characterize ridges, we constructed a new HTM based on the draft human genome sequence (HTMseq). Expression of 25,003 genes can be analyzed online in a multitude of tissues (http://bioinfo.amc.uva.nl/HTMseq). Ridges are found to be very gene-dense domains with a high GC content, a high SINE repeat density, and a low LINE repeat density. Genes in ridges have significantly shorter introns than genes outside of ridges. The HTMseq also identifies a significant clustering of weakly expressed genes in domains with fully opposite characteristics (antiridges). Both types of domains are open to tissue-specific expression regulation, but the maximal expression levels in ridges are considerably higher than in antiridges. Ridges are therefore an integral part of a higher order structure in the genome related to transcriptional regulation.

Base Composition↗

REDUCE: An online tool for inferring cis-regulatory elements and transcriptional module activities from microarray data.

REDUCE is a motif-based regression method for microarray analysis. The only required inputs are (i) a single genome-wide set of absolute or relative mRNA abundances and (ii) the DNA sequence of the regulatory region associated with each gene that is probed. Currently supported organisms are yeast, worm and fly; it is an open question whether in its current incarnation our approach can be used for mouse or human. REDUCE uses unbiased statistics to identify oligonucleotide motifs whose occurrence in the regulatory region of a gene correlates with the level of mRNA expression. Regression analysis is used to infer the activity of the transcriptional module associated with each motif. REDUCE is available online at http://bussemaker.bio.columbia.edu/reduce/. This web site provides functionality for the upload and management of microarray data. REDUCE analysis results can be viewed and downloaded, and optionally be shared with other users or made publicly accessible.

Algorithms↗

Revisiting the codon adaptation index from a whole-genome perspective: analyzing the relationship between gene expression and codon occurrence in yeast using a variety of models.

Highly expressed genes in many bacteria and small eukaryotes often have a strong compositional bias, in terms of codon usage. Two widely used numerical indices, the codon adaptation index (CAI) and the codon usage, use this bias to predict the expression level of genes. When these indices were first introduced, they were based on fairly simple assumptions about which genes are most highly expressed: the CAI was originally based on the codon composition of a set of only 24 highly expressed genes, and the codon usage on assumptions about which functional classes of genes are highly expressed in fast-growing bacteria. Given the recent advent of genome-wide expression data, we should be able to improve on these assumptions. Here, we measure, in yeast, the degree to which consideration of the current genome-wide expression data sets improves the performance of both numerical indices. Indeed, we find that by changing the parameterization of each model its correlation with actual expression levels can be somewhat improved, although both indices are fairly insensitive to the exact way they are parameterized. This insensitivity indicates a consistent codon bias amongst highly expressed genes. We also attempt direct linear regression of codon composition against genome-wide expression levels (and protein abundance data). This has some similarity with the CAI formalism and yields an alternative model for the prediction of expression levels based on the coding sequences of genes. More information is available at http://bioinfo.mbb.yale.edu/expression/codons.

Codon↗

Genomic binding by the Drosophila Myc, Max, Mad/Mnt transcription factor network.

The Myc/Max/Mad transcription factor network is critically involved in cell behavior; however, there is relatively little information on its genomic binding sites. We have employed the DamID method to carry out global genomic mapping of the Drosophila Myc, Max, and Mad/Mnt proteins. Each protein was tethered to Escherichia coli DNA adenine-methyltransferase (Dam) permitting methylation proximal to in vivo binding sites in Kc cells. Microarray analyses of methylated DNA fragments reveals binding to multiple loci on all major Drosophila chromosomes. This approach also reveals dynamic interactions among network members as we find that increased levels of dMax influence the extent of dMyc, but not dMnt, binding. Computer analysis using the REDUCE algorithm demonstrates that binding regions correlate with the presence of E-boxes, CG repeats, and other sequence motifs. The surprisingly large number of directly bound loci ( approximately 15% of coding regions) suggests that the network interacts widely with the genome. Furthermore, we employ microarray expression analysis to demonstrate that hundreds of DamID-binding loci correspond to genes whose expression is directly regulated by dMyc in larvae. These results suggest that a fundamental aspect of Max network function involves widespread binding and regulation of gene expression.

Animals↗

Genomewide analysis of Drosophila GAGA factor target genes reveals context-dependent DNA binding.

The association of sequence-specific DNA-binding factors with their cognate target sequences in vivo depends on the local molecular context, yet this context is poorly understood. To address this issue, we have performed genomewide mapping of in vivo target genes of Drosophila GAGA factor (GAF). The resulting list of approximately 250 target genes indicates that GAF regulates many cellular pathways. We applied unbiased motif-based regression analysis to identify the sequence context that determines GAF binding. Our results confirm that GAF selectively associates with (GA)(n) repeat elements in vivo. GAF binding occurs in upstream regulatory regions, but less in downstream regions. Surprisingly, GAF binds abundantly to introns but is virtually absent from exons, even though the density of (GA)(n) is roughly the same. Intron binding occurs equally frequently in last introns compared with first introns, suggesting that GAF may not only regulate transcription initiation, but possibly also elongation. We provide evidence for cooperative binding of GAF to closely spaced (GA)(n) elements and explain the lack of GAF binding to exons by the absence of such closely spaced GA repeats. Our approach for revealing determinants of context-dependent DNA binding will be applicable to many other transcription factors.

Amino Acid Motifs↗

Hap4p overexpression in glucose-grown Saccharomyces cerevisiae induces cells to enter a novel metabolic state.

BACKGROUND: Metabolic and regulatory gene networks generally tend to be stable. However, we have recently shown that overexpression of the transcriptional activator Hap4p in yeast causes cells to move to a state characterized by increased respiratory activity. To understand why overexpression of HAP4 is able to override the signals that normally result in glucose repression of mitochondrial function, we analyzed in detail the changes that occur in these cells. RESULTS: Whole-genome expression profiling and fingerprinting of the regulatory activity network show that HAP4 overexpression provokes changes that also occur during the diauxic shift. Overexpression of HAP4, however, primarily acts on mitochondrial function and biogenesis. In fact, a number of nuclear genes encoding mitochondrial proteins are induced to a greater extent than in cells that have passed through a normal diauxic shift: in addition to genes required for mitochondrial energy conservation they include genes encoding mitochondrial ribosomal proteins. CONCLUSIONS: We show that overproduction of a single nuclear transcription factor enables cells to move to a novel state that displays features typical of, but clearly not identical to, other derepressed states.

CCAAT-Binding Factor↗

Identification of genes expressed in C. elegans touch receptor neurons.

The extent of gene regulation in cell differentiation is poorly understood. We previously used saturation mutagenesis to identify 18 genes that are needed for the development and function of a single type of sensory neuron--the touch receptor neuron for gentle touch in Caenorhabditis elegans. One of these genes, mec-3, encodes a transcription factor that controls touch receptor differentiation. By culturing and isolating wild-type and mec-3 mutant cells from embryos and applying their amplified RNA to DNA microarrays, here we have identified genes that are known to be expressed in touch receptors, a previously uncloned gene (mec-17) that is needed for maintaining touch receptor differentiation, and more than 50 previously unknown mec-3-dependent genes. These genes are randomly distributed in the genome and under-represented both for genes that are co-expressed in operons and for multiple members of gene families. Using regions 5' of the start codon of the first 20 genes, we have also identified an over-represented heptanucleotide, AATGCAT, that is needed for the expression of touch receptor genes.

Amino Acid Sequence↗

Dissection of transient oxidative stress response in Saccharomyces cerevisiae by using DNA microarrays.

Yeast cells were grown in glucose-limited chemostat cultures and forced to switch to a new carbon source, the fatty acid oleate. Alterations in gene expression were monitored using DNA microarrays combined with bioinformatics tools, among which was included the recently developed algorithm REDUCE. Immediately after the switch to oleate, a transient and very specific stress response was observed, followed by the up-regulation of genes encoding peroxisomal enzymes required for fatty acid metabolism. The stress response included up-regulation of genes coding for enzymes to keep thioredoxin and glutathione reduced, as well as enzymes required for the detoxification of reactive oxygen species. Among the genes coding for various isoenzymes involved in these processes, only a specific subset was expressed. Not the general stress transcription factors Msn2 and Msn4, but rather the specific factor Yap1p seemed to be the main regulator of the stress response. We ascribe the initiation of the oxidative stress response to a combination of poor redox flux and fatty acid-induced uncoupling of the respiratory chain during the metabolic reprogramming phase.

Active Transport, Cell Nucleus↗