PubMed Health⌕ Search

Biomedical subjects

Matthew W Hahn

Publications and source records attributed to Matthew W Hahn.

At least 19 recordsLinked to original sources

The evolution of mammalian gene families.

Gene families are groups of homologous genes that are likely to have highly similar functions. Differences in family size due to lineage-specific gene duplication and gene loss may provide clues to the evolutionary forces that have shaped mammalian genomes. Here we analyze the gene families contained within the whole genomes of human, chimpanzee, mouse, rat, and dog. In total we find that more than half of the 9,990 families present in the mammalian common ancestor have either expanded or contracted along at least one lineage. Additionally, we find that a large number of families are completely lost from one or more mammalian genomes, and a similar number of gene families have arisen subsequent to the mammalian common ancestor. Along the lineage leading to modern humans we infer the gain of 689 genes and the loss of 86 genes since the split from chimpanzees, including changes likely driven by adaptive natural selection. Our results imply that humans and chimpanzees differ by at least 6% (1,418 of 22,000 genes) in their complement of genes, which stands in stark contrast to the oft-cited 1.5% difference between orthologous nucleotide sequences. This genomic "revolving door" of gene gain and loss represents a large number of genetic differences separating humans from our closest relatives.

Animals↗

Detecting natural selection on cis-regulatory DNA.

Changes in transcriptional regulation play an important role in the genetic basis for evolutionary change. Here I review a growing body of literature that seeks to determine the forces governing the non-coding regulatory sequences underlying these changes. I address the challenges present in studying natural selection without the familiar structure and regularity of protein-coding sequences, but show that most tests of neutrality that have been used for coding regions are applicable to non-coding regions, albeit with some caveats. While some experimental investment is necessary to identify heritable regulatory variation, the most basic inferences about selection require very little functional information. A growing body of research on cis-regulatory variation has uncovered all the forms of selection common to coding regions, in addition to novel forms of selection. An emerging pattern seems to be the ubiquity of local adaptation and balancing selection, possibly due to the greater freedom organisms have to fine-tune gene expression without changing protein function. It is clear from multiple single locus and whole genome studies of non-coding regulatory DNA that the effects of natural selection reach far beyond the start and stop codons.

Animals↗

CAFE: a computational tool for the study of gene family evolution.

SUMMARY: We present CAFE (Computational Analysis of gene Family Evolution), a tool for the statistical analysis of the evolution of the size of gene families. It uses a stochastic birth and death process to model the evolution of gene family sizes over a phylogeny. For a specified phylogenetic tree, and given the gene family sizes in the extant species, CAFE can estimate the global birth and death rate of gene families, infer the most likely gene family size at all internal nodes, identify gene families that have accelerated rates of gain and loss (quantified by a p-value) and identify which branches cause the p-value to be small for significant families. AVAILABILITY: Software is available from http://www.bio.indiana.edu/~hahnlab/Software.html

Algorithms↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Accurate inference and estimation in population genomics.

Both intra- and interspecific genomic comparisons have revealed local similarities in the level and frequency of mutational variation, as well as in patterns of gene expression. This autocorrelation between measurements leads to violations of assumptions of independence in many statistical methods, resulting in misleading and incorrect inferences. Here I show that autocorrelation can be due to many factors and is present across the genome. Using a one-dimensional spatial stochastic model, I further show how previous results can be employed to correct for autocorrelation along chromosomes in population and comparative genomics research. When multiple hypothesis tests are autocorrelated, I demonstrate that a simple correction can lead to increased power in statistical inference. I present a preliminary analysis of population genomic data from Drosophila simulans to show the ubiquity of autocorrelation and applicability of the methods proposed here.

Animals↗

Genomic islands of speciation in Anopheles gambiae.

The African malaria mosquito, Anopheles gambiae sensu stricto (A. gambiae), provides a unique opportunity to study the evolution of reproductive isolation because it is divided into two sympatric, partially isolated subtaxa known as M form and S form. With the annotated genome of this species now available, high-throughput techniques can be applied to locate and characterize the genomic regions contributing to reproductive isolation. In order to quantify patterns of differentiation within A. gambiae, we hybridized population samples of genomic DNA from each form to Affymetrix GeneChip microarrays. We found that three regions, together encompassing less than 2.8 Mb, are the only locations where the M and S forms are significantly differentiated. Two of these regions are adjacent to centromeres, on Chromosomes 2L and X, and contain 50 and 12 predicted genes, respectively. Sequenced loci in these regions contain fixed differences between forms and no shared polymorphisms, while no fixed differences were found at nearby control loci. The third region, on Chromosome 2R, contains only five predicted genes; fixed differences in this region were also verified by direct sequencing. These "speciation islands" remain differentiated despite considerable gene flow, and are therefore expected to contain the genes responsible for reproductive isolation. Much effort has recently been applied to locating the genes and genetic changes responsible for reproductive isolation between species. Though much can be inferred about speciation by studying taxa that have diverged for millions of years, studying differentiation between taxa that are in the early stages of isolation will lead to a clearer view of the number and size of regions involved in the genetics of speciation. Despite appreciable levels of gene flow between the M and S forms of A. gambiae, we were able to isolate three small regions of differentiation where genes responsible for ecological and behavioral isolation are likely to be located. We expect reproductive isolation to be due to changes at a small number of loci, as these regions together contain only 67 predicted genes. Concentrating future mapping experiments on these regions should reveal the genes responsible for reproductive isolation between forms.

Animals↗

Evolutionary genomics: codon bias and selection on single genomes.

The idea that natural selection on genes might be detected using only a single genome has been put forward by Plotkin and colleagues, who present a method that they claim can detect selection without the need for comparative data and which, if correct, would confer greater power of analysis with less information. Here we argue that their method depends on assumptions that confound their conclusions and that, even if these assumptions were valid, the authors' inferences about adaptive natural selection are unjustified.

Bias↗

Estimating the tempo and mode of gene family evolution from comparative genomic data.

Comparison of whole genomes has revealed that changes in the size of gene families among organisms is quite common. However, there are as yet no models of gene family evolution that make it possible to estimate ancestral states or to infer upon which lineages gene families have contracted or expanded. In addition, large differences in family size have generally been attributed to the effects of natural selection, without a strong statistical basis for these conclusions. Here we use a model of stochastic birth and death for gene family evolution and show that it can be efficiently applied to multispecies genome comparisons. This model takes into account the lengths of branches on phylogenetic trees, as well as duplication and deletion rates, and hence provides expectations for divergence in gene family size among lineages. The model offers both the opportunity to identify large-scale patterns in genome evolution and the ability to make stronger inferences regarding the role of natural selection in gene family expansion or contraction. We apply our method to data from the genomes of five yeast species to show its applicability.

Evolution, Molecular↗

Ancient and recent positive selection transformed opioid cis-regulation in humans.

Changes in the cis-regulation of neural genes likely contributed to the evolution of our species' unique attributes, but evidence of a role for natural selection has been lacking. We found that positive natural selection altered the cis-regulation of human prodynorphin, the precursor molecule for a suite of endogenous opioids and neuropeptides with critical roles in regulating perception, behavior, and memory. Independent lines of phylogenetic and population genetic evidence support a history of selective sweeps driving the evolution of the human prodynorphin promoter. In experimental assays of chimpanzee-human hybrid promoters, the selected sequence increases transcriptional inducibility. The evidence for a change in the response of the brain's natural opioids to inductive stimuli points to potential human-specific characteristics favored during evolution. In addition, the pattern of linked nucleotide and microsatellite variation among and within modern human populations suggests that recent selection, subsequent to the fixation of the human-specific mutations and the peopling of the globe, has favored different prodynorphin cis-regulatory alleles in different parts of the world.

Alleles↗

Comparative genomics of centrality and essentiality in three eukaryotic protein-interaction networks.

Most proteins do not evolve in isolation, but as components of complex genetic networks. Therefore, a protein's position in a network may indicate how central it is to cellular function and, hence, how constrained it is evolutionarily. To look for an effect of position on evolutionary rate, we examined the protein-protein interaction networks in three eukaryotes: yeast, worm, and fly. We find that the three networks have remarkably similar structure, such that the number of interactors per protein and the centrality of proteins in the networks have similar distributions. Proteins that have a more central position in all three networks, regardless of the number of direct interactors, evolve more slowly and are more likely to be essential for survival. Our results are thus consistent with a classic proposal of Fisher's that pleiotropy constrains evolution.

Biological Evolution↗

Disentangling the effects of demography and selection in human history.

Demographic events affect all genes in a genome, whereas natural selection has only local effects. Using publicly available data from 151 loci sequenced in both European-American and African-American populations, we attempt to distinguish the effects of demography and selection. To analyze large sets of population genetic data such as this one, we introduce "Perlymorphism," a Unix-based suite of analysis tools. Our analyses show that the demographic histories of human populations can account for a large proportion of effects on the level and frequency of variation across the genome. The African-American population shows both a higher level of nucleotide diversity and more negative values of Tajima's D statistic than does a European-American population. Using coalescent simulations, we show that the significantly negative values of the D statistic in African-Americans and the positive values in European-Americans are well explained by relatively simple models of population admixture and bottleneck, respectively. Working within these nonequilibrium frameworks, we are still able to show deviations from neutral expectations at a number of loci, including ABO and TRPV6. In addition, we show that the frequency spectrum of mutations--corrected for levels of polymorphism--is correlated with recombination rate only in European-Americans. These results are consistent with repeated selective sweeps in non-African populations, in agreement with recent reports using microsatellite data.

ABO Blood-Group System↗

Positive selection on MMP3 regulation has shaped heart disease risk.

BACKGROUND: The evolutionary forces of mutation, natural selection, and genetic drift shape the pattern of phenotypic variation in nature, but the roles of these forces in defining the distributions of particular traits have been hard to disentangle. To better understand the mechanisms contributing to common variation in humans, we investigated the evolutionary history of a functional polymorphism in the upstream regulatory region of the MMP3 gene. This single base pair insertion/deletion variant, which results in a run of either 5 or 6 thymidines 1608 bp from the transcription start site, alters transcription factor binding and influences levels of MMP3 mRNA and protein. The polymorphism contributes to variation in arterial traits and to the risk of coronary heart disease and its progression. RESULTS: Phylogenetic and population genetic analysis of primate sequences indicate that the binding site region is rapidly evolving and has been a hot spot for mutation for tens of millions of years. We also find evidence for the action of positive selection, beginning approximately 24,000 years ago, increasing the frequency of the high-expression allele in Europe but not elsewhere. Positive selection is evident in statistical tests of differentiation among populations and haplotype diversity within populations. Europeans have greater arterial elasticity and suffer dramatically fewer coronary heart disease events than they would have had this selection not occurred. CONCLUSIONS: Locally elevated mutation rates and strong positive selection on a cis-regulatory variant have shaped contemporary phenotypic variation and public health.

Animals↗

Random drift and large shifts in popularity of dog breeds.

A simple model of random copying among individuals, similar to the population genetic model of random drift, can predict the variability in the popularity of cultural variants. Here, we show that random drift also explains a biologically relevant cultural phenomenon--changes in the distributions of popularity of dog breeds in the United States in each of the past 50 years. There are, however, interesting deviations from the model that involve large changes in the popularity of certain breeds. By identifying meaningful departures from our null model, we show how it can serve as a foundation for studying culture change quantitatively, using the tools of population genetics.

Animals↗

Random drift and culture change.

We show that the frequency distributions of cultural variants, in three different real-world examples--first names, archaeological pottery and applications for technology patents--follow power laws that can be explained by a simple model of random drift. We conclude that cultural and economic choices often reflect a decision process that is value-neutral; this result has far-reaching testable implications for social-science research.

Computer Simulation↗

Molecular evolution in large genetic networks: does connectivity equal constraint?

Genetic networks show a broad-tailed distribution of the number of interaction partners per protein, which is consistent with a power-law. It has been proposed that such broad-tailed distributions are observed because they confer robustness against mutations to the network. We evaluate this hypothesis for two genetic networks, that of the E. coli core intermediary metabolism and that of the yeast protein-interaction network. Specifically, we test the hypothesis through one of its key predictions: highly connected proteins should be more important to the cell and, thus, subject to more severe selective and evolutionary constraints. We find, however, that no correlation between highly connected proteins and evolutionary rate exists in the E. coli metabolic network and that there is only a weak correlation in the yeast protein-interaction network. Furthermore, we show that the observed correlation is function-specific within the protein-interaction network: only genes involved in the cell cycle and transcription show significant correlations. Our work sheds light on conflicting results by previous researchers by comparing data from multiple types of protein-interaction datasets and by using a closely related species as a reference taxon. The finding that highly connected proteins can tolerate just as many amino acid substitutions as other proteins leads us to conclude that power-laws in cellular networks do not reflect selection for mutational robustness.

Energy Metabolism↗

Population genetic and phylogenetic evidence for positive selection on regulatory mutations at the factor VII locus in humans.

The abundance of cis-regulatory polymorphisms in humans suggests that many may have been important in human evolution, but evidence for their role is relatively rare. Four common polymorphisms in the 5' promoter region of factor VII (F7), a coagulation factor, have been shown to affect its transcription and protein abundance both in vitro and in vivo. Three of these polymorphisms have low-frequency alleles that decrease expression of F7 and may provide protection against myocardial infarction (heart attacks). The fourth polymorphism has a minor allele that increases the level of transcription. To look for evidence of natural selection on the cis-regulatory variants flanking F7, we genotyped three of the polymorphisms in six Old World populations for which we also have data from a group of putatively neutral SNPs. Our population genetic analysis shows evidence for selection within humans; surprisingly, the strongest evidence is due to a large increase in frequency of the high-expression variant in Singaporean Chinese. Further characterization of a Japanese population shows that at least part of the increase in frequency of the high-expression allele is found in other East Asian populations. In addition, to examine interspecific patterns of selection we sequenced the homologous 5' noncoding region in chimpanzees, bonobos, a gorilla, an orangutan, and a baboon. Analysis of these data reveals an excess of fixed differences within transcription factor binding sites along the human lineage. Our results thus further support the hypothesis that regulatory mutations have been important in human evolution.

Animals↗

Positive selection on a human-specific transcription factor binding site regulating IL4 expression.

A single nucleotide polymorphism in the promoter of the multifunctional cytokine Interleukin 4 (IL4) affects the binding of NFAT, a key transcriptional activator of IL4 in T cells. This regulatory polymorphism influences the balance of cytokine signaling in the immune system, with important consequences-positive and negative-for human health. We determined that the NFAT binding site is unique to humans; it arose by point mutation along the lineage separating humans from other great apes. We show that its frequency distribution among human subpopulations has been shaped by the balance of selective forces on IL4's diverse roles. New statistical approaches, based on parametric and nonparametric comparisons to neutral variants typed in the same individuals, indicate that differentiation among subpopulations at the IL4 promoter polymorphism is too great to be attributed to neutral drift. The allele frequencies of this binding site represent local adaptation to diverse pathogenic challenges; disease states associated with the common derived allele are side-effects of positive selection on other IL4 functions.

Animals↗