PubMed Health⌕ Search

Biomedical subjects

Kevin P White

Publications and source records attributed to Kevin P White.

At least 19 recordsLinked to original sources

Hotspots of transcription factor colocalization in the genome of Drosophila melanogaster.

Regulation of gene expression is a highly complex process that requires the concerted action of many proteins, including sequence-specific transcription factors, cofactors, and chromatin proteins. In higher eukaryotes, the interplay between these proteins and their interactions with the genome still is poorly understood. We systematically mapped the in vivo binding sites of seven transcription factors with diverse physiological functions, five cofactors, and two heterochromatin proteins at approximately 1-kb resolution in a 2.9 Mb region of the Drosophila melanogaster genome. Surprisingly, all tested transcription factors and cofactors show strongly overlapping localization patterns, and the genome contains many "hotspots" that are targeted by all of these proteins. Several control experiments show that the strong overlap is not an artifact of the techniques used. Colocalization hotspots are 1-5 kb in size, spaced on average by approximately 50 kb, and preferentially located in regions of active transcription. We provide evidence that protein-protein interactions play a role in the hotspot association of some transcription factors. Colocalization hotspots constitute a previously uncharacterized type of feature in the genome of Drosophila, and our results provide insights into the general targeting mechanisms of transcription regulators in a higher eukaryote.

Animals↗

Chromosomal distribution of PcG proteins during Drosophila development.

Polycomb group (PcG) proteins are able to maintain the memory of silent transcriptional states of homeotic genes throughout development. In Drosophila, they form multimeric complexes that bind to specific DNA regulatory elements named PcG response elements (PREs). To date, few PREs have been identified and the chromosomal distribution of PcG proteins during development is unknown. We used chromatin immunoprecipitation (ChIP) with genomic tiling path microarrays to analyze the binding profile of the PcG proteins Polycomb (PC) and Polyhomeotic (PH) across 10 Mb of euchromatin. We also analyzed the distribution of GAGA factor (GAF), a sequence-specific DNA binding protein that is found at most previously identified PREs. Our data show that PC and PH often bind to clustered regions within large loci that encode transcription factors which play multiple roles in developmental patterning and in the regulation of cell proliferation. GAF co-localizes with PC and PH to a limited extent, suggesting that GAF is not a necessary component of chromatin at PREs. Finally, the chromosome-association profile of PC and PH changes during development, suggesting that the function of these proteins in the regulation of some of their target genes might be more dynamic than previously anticipated.

Animals↗

Microparadigms: chains of collective reasoning in publications about molecular interactions.

We analyzed a very large set of molecular interactions that had been derived automatically from biological texts. We found that published statements, regardless of their verity, tend to interfere with interpretation of the subsequent experiments and, therefore, can act as scientific "microparadigms," similar to dominant scientific theories [Kuhn, T. S. (1996) The Structure of Scientific Revolutions (Univ. Chicago Press, Chicago)]. Using statistical tools, we measured the strength of the influence of a single published statement on subsequent interpretations. We call these measured values the momentums of the published statements and treat separately the majority and minority of conflicting statements about the same molecular event. Our results indicate that, when building biological models based on published experimental data, we may have to treat the data as highly dependent-ordered sequences of statements (i.e., chains of collective reasoning) rather than unordered and independent experimental observations. Furthermore, our computations indicate that our data set can be interpreted in two very different ways (two "alternative universes"): one is an "optimists' universe" with a very low incidence of false results (<5%), and another is a "pessimists' universe" with an extraordinarily high rate of false results (>90%). Our computations deem highly unlikely any milder intermediate explanation between these two extremes.

Computer Simulation↗

Expression profiling in primates reveals a rapid evolution of human transcription factors.

Although it has been hypothesized for thirty years that many human adaptations are likely to be due to changes in gene regulation, almost nothing is known about the modes of natural selection acting on regulation in primates. Here we identify a set of genes for which expression is evolving under natural selection. We use a new multi-species complementary DNA array to compare steady-state messenger RNA levels in liver tissues within and between humans, chimpanzees, orangutans and rhesus macaques. Using estimates from a linear mixed model, we identify a set of genes for which expression levels have remained constant across the entire phylogeny (approximately 70 million years), and are therefore likely to be under stabilizing selection. Among the top candidates are five genes with expression levels that have previously been shown to be altered in liver carcinoma. We also find a number of genes with similar expression levels among non-human primates but significantly elevated or reduced expression in the human lineage, features that point to the action of directional selection. Among the gene set with a human-specific increase in expression, there is an excess of transcription factors; the same is not true for genes with increased expression in chimpanzee.

Animals↗

Detecting transcriptionally active regions using genomic tiling arrays.

We have developed a method for interpreting genomic tiling array data, implemented as the program TranscriptionDetector. Probed loci expressed above background are identified by combining replicates in a way that makes minimal assumptions about the data. We performed medium-resolution Anopheles gambiae tiling array experiments and found extensive transcription of both coding and non-coding regions. Our method also showed improved detection of transcriptional units when applied to high-density tiling array data for ten human chromosomes.

Animals↗

A mutation accumulation assay reveals a broad capacity for rapid evolution of gene expression.

Mutation is the ultimate source of biological diversity because it generates the variation that fuels evolution. Gene expression is the first step by which an organism translates genetic information into developmental change. Here we estimate the rate at which mutation produces new variation in gene expression by measuring transcript abundances across the genome during the onset of metamorphosis in 12 initially identical Drosophila melanogaster lines that independently accumulated mutations for 200 generations. We find statistically significant mutational variation for 39% of the genome and a wide range of variability across corresponding genes. As genes are upregulated in development their variability decreases, and as they are downregulated it increases, indicating that developmental context affects the evolution of gene expression. A strong correlation between mutational variance and environmental variance shows that there is the potential for widespread canalization. By comparing the evolutionary rates that we report here with differences between species, we conclude that gene expression does not evolve according to strictly neutral models. Although spontaneous mutations have the potential to generate abundant variation in gene expression, natural variation is relatively constrained.

Animals↗

Immune signaling pathways regulating bacterial and malaria parasite infection of the mosquito Anopheles gambiae.

We show that, in the malaria vector Anopheles gambiae, expression of Cecropin 1 is regulated by REL2, an NF-kappaB-like transcription factor orthologous to Drosophila Relish. Through alternative splicing, REL2 produces a full-length (REL2-F) and a shorter (REL2-S) protein isoform lacking the inhibitory ankyrin repeats and death domain. RNA interference experiments show that, in contrast to Drosophila Relish, which responds solely to Gram-negative bacteria, the Anopheles REL2-F and REL2-S isoforms are involved in defense against the Gram-positive Staphylococcus aureus and the Gram-negative Escherichia coli bacteria, respectively. REL2-F also regulates the intensity of mosquito infection with the malaria parasite, Plasmodium berghei. The adaptor IMD shares the same activities as REL2-F. Microarray analysis identified 10 additional genes regulated by REL2, including CEC3, GAM1, and LRIM1.

Alternative Splicing↗

Comparative genome sequencing of Drosophila pseudoobscura: chromosomal, gene, and cis-element evolution.

We have sequenced the genome of a second Drosophila species, Drosophila pseudoobscura, and compared this to the genome sequence of Drosophila melanogaster, a primary model organism. Throughout evolution the vast majority of Drosophila genes have remained on the same chromosome arm, but within each arm gene order has been extensively reshuffled, leading to a minimum of 921 syntenic blocks shared between the species. A repetitive sequence is found in the D. pseudoobscura genome at many junctions between adjacent syntenic blocks. Analysis of this novel repetitive element family suggests that recombination between offset elements may have given rise to many paracentric inversions, thereby contributing to the shuffling of gene order in the D. pseudoobscura lineage. Based on sequence similarity and synteny, 10,516 putative orthologs have been identified as a core gene set conserved over 25-55 million years (Myr) since the pseudoobscura/melanogaster divergence. Genes expressed in the testes had higher amino acid sequence divergence than the genome-wide average, consistent with the rapid evolution of sex-specific proteins. Cis-regulatory sequences are more conserved than random and nearby sequences between the species--but the difference is slight, suggesting that the evolution of cis-regulatory elements is flexible. Overall, a pattern of repeat-mediated chromosomal rearrangement, and high coadaptation of both male genes and cis-regulatory sequences emerges as important themes of genome divergence between these species of Drosophila.

Animals↗

Multi-species microarrays reveal the effect of sequence divergence on gene expression profiles.

Interspecies comparisons of gene expression levels will increase our understanding of the evolution of transcriptional mechanisms and help to identify targets of natural selection. This approach holds particular promise for apes, as many human-specific adaptations are thought to result from differences in gene expression rather than in coding sequence. To date, however, all studies directly comparing interspecies gene expression have been performed on single-species arrays, so that it has been impossible to distinguish differential hybridization due to sequence mismatches from underlying expression differences. To evaluate the severity of this potential problem, we constructed a new multiprimate cDNA array using probes from human, chimpanzee, orangutan, and rhesus. We find a large effect of sequence divergence on hybridization signal, even in the closest pair of species, human and chimpanzee. By comparing single-species array analyses with results from multispecies arrays, we examine how estimates of differential gene expression are affected by sequence divergence. Our results indicate that naive use of single-species arrays in direct interspecies comparisons can yield spurious results.

Animals↗

A gene expression map for the euchromatic genome of Drosophila melanogaster.

We used a maskless photolithography method to produce DNA oligonucleotide microarrays with unique probe sequences tiled throughout the genome of Drosophila melanogaster and across predicted splice junctions. RNA expression of protein coding and nonprotein coding sequences was determined for each major stage of the life cycle, including adult males and females. We detected transcriptional activity for 93% of annotated genes and RNA expression for 41% of the probes in intronic and intergenic sequences. Comparison to genome-wide RNA interference data and to gene annotations revealed distinguishable levels of expression for different classes of genes and higher levels of expression for genes with essential cellular functions. Differential splicing was observed in about 40% of predicted genes, and 5440 previously unknown splice forms were detected. Genes within conserved regions of synteny with D. pseudoobscura had highly correlated expression; these regions ranged in length from 10 to 900 kilobase pairs. The expressed intergenic and intronic sequences are more likely to be evolutionarily conserved than nonexpressed ones, and about 15% of them appear to be developmentally regulated. Our results provide a draft expression map for the entire nonrepetitive genome, which reveals a much more extensive and diverse set of expressed sequences than was previously predicted.

Algorithms↗

Duplicate genes increase gene expression diversity within and between species.

Using microarray gene expression data from several Drosophila species and strains, we show that duplicated genes, compared with single-copy genes, significantly increase gene expression diversity during development. We show further that duplicate genes tend to cause expression divergences between Drosophila species (or strains) to evolve faster than do single-copy genes. This conclusion is also supported by data from different yeast strains.

Animals↗

Probabilistic inference of molecular networks from noisy data sources.

Information on molecular networks, such as networks of interacting proteins, comes from diverse sources that contain remarkable differences in distribution and quantity of errors. Here, we introduce a probabilistic model useful for predicting protein interactions from heterogeneous data sources. The model describes stochastic generation of protein-protein interaction networks with real-world properties, as well as generation of two heterogeneous sources of protein-interaction information: research results automatically extracted from the literature and yeast two-hybrid experiments. Based on the domain composition of proteins, we use the model to predict protein interactions for pairs of proteins for which no experimental data are available. We further explore the prediction limits, given experimental data that cover only part of the underlying protein networks. This approach can be extended naturally to include other types of biological data sources.

Algorithms↗

Protein-DNA interaction mapping using genomic tiling path microarrays in Drosophila.

We demonstrate the use of a chromosomal walk (or "tiling path") printed as DNA microarrays for mapping protein-DNA interactions across large regions of contiguous genomic DNA in Drosophila melanogaster. Microarrays were constructed with genomic DNA fragments 430-920 bp in length, covering 2.9 million base pairs of the Adh-cactus region of chromosome 2 and 85,000 base pairs of the 82F region of chromosome 3. We performed DNA localization mapping for the heterochromatin protein HP1 and for the sequence-specific GAGA transcription factor, producing a comprehensive, high-resolution map of in vivo protein-DNA interactions throughout these regions of the Drosophila genome.

Amino Acid Motifs↗

Analysis of the eye developmental pathway in Drosophila using DNA microarrays.

Pax-6 genes encode evolutionarily conserved transcription factors capable of activating the gene-expression program required to build an eye. When ectopically expressed in Drosophila imaginal discs, Pax-6 genes induce the eye formation on the corresponding appendages of the adult fly. We used two different Drosophila full-genome DNA microarrays to compare gene expression in wild-type leg discs versus leg discs where eyeless, one of the two Drosophila Pax-6 genes, was ectopically expressed. We validated these data by analyzing the endogenous expression of selected genes in eye discs and identified 371 genes that are expressed in the eye imaginal discs and up-regulated when an eye morphogenetic field is ectopically induced in the leg discs. These genes mainly encode transcription factors involved in photoreceptor specification, signal transducers, cell adhesion molecules, and proteins involved in cell division. As expected, genes already known to act downstream of eyeless during eye development were identified, together with a group of genes that were not yet associated with eye formation.

Animals↗