PubMed Health⌕ Search

Biomedical subjects

Jeffrey D Wall

Publications and source records attributed to Jeffrey D Wall.

15 recordsLinked to original sources

A worldwide survey of haplotype variation and linkage disequilibrium in the human genome.

Recent genomic surveys have produced high-resolution haplotype information, but only in a small number of human populations. We report haplotype structure across 12 Mb of DNA sequence in 927 individuals representing 52 populations. The geographic distribution of haplotypes reflects human history, with a loss of haplotype diversity as distance increases from Africa. Although the extent of linkage disequilibrium (LD) varies markedly across populations, considerable sharing of haplotype structure exists, and inferred recombination hotspot locations generally match across groups. The four samples in the International HapMap Project contain the majority of common haplotypes found in most populations: averaging across populations, 83% of common 20-kb haplotypes in a population are also common in the most similar HapMap sample. Consequently, although the portability of tag SNPs based on the HapMap is reduced in low-LD Africans, the HapMap will be helpful for the design of genome-wide association mapping studies in nearly all human populations.

Chromosome Mapping↗

Archaic admixture in the human genome.

One of the enduring questions in the evolution of our species surrounds the fate of 'archaic' forms of Homo. Did Neanderthals go extinct without interbreeding with modern humans 25-40 thousand years ago or are their genes present among modern-day Europeans? Recent work suggests that Neanderthals and an as yet unidentified archaic African population contributed to at least 5% of the modern European and West African gene pools, respectively. Extensive sequencing of Neanderthal and other archaic human nuclear DNA has the potential to answer this question definitively within the next few years.

Animals↗

Estimating recombination rates from single-nucleotide polymorphisms using summary statistics.

We describe a novel method for jointly estimating crossing-over and gene-conversion rates from population genetic data using summary statistics. The performance of our method was tested on simulated data sets and compared with the composite-likelihood method of R. R. Hudson. For several realistic parameter values, the new method performed similarly to the composite-likelihood approach for estimating crossing-over rates and better when estimating gene-conversion rates. We used our method to analyze a human data set recently genotyped by Perlegen Sciences.

Computer Simulation↗

Possible ancestral structure in human populations.

Determining the evolutionary relationships between fossil hominid groups such as Neanderthals and modern humans has been a question of enduring interest in human evolutionary genetics. Here we present a new method for addressing whether archaic human groups contributed to the modern gene pool (called ancient admixture), using the patterns of variation in contemporary human populations. Our method improves on previous work by explicitly accounting for recent population history before performing the analyses. Using sequence data from the Environmental Genome Project, we find strong evidence for ancient admixture in both a European and a West African population (p approximately 10(-7)), with contributions to the modern gene pool of at least 5%. While Neanderthals form an obvious archaic source population candidate in Europe, there is not yet a clear source population candidate in West Africa.

Animals↗

Relative influences of crossing over and gene conversion on the pattern of linkage disequilibrium in Arabidopsis thaliana.

In this article we infer the rates of gene conversion and crossing over in Arabidopsis thaliana from population genetic data. Our data set is a genomewide survey consisting of 1347 fragments of length 600 bp sequenced in 96 accessions. It has several orders of magnitude more markers than any previous nonhuman study. This allows for more accurate inference as well as a detailed comparison between theoretical expectations and observations. Our methodology is specifically set to account for deviations such as recurrent mutations or a skewed frequency spectrum. We found that even if some components of the model clearly do not fit, the pattern of LD conforms to theoretical expectations quite well. The ratio of gene conversion to crossing over is estimated to be around one. We also find evidence for fine-scale variations of the crossing-over rate.

Arabidopsis↗

The pattern of polymorphism in Arabidopsis thaliana.

We resequenced 876 short fragments in a sample of 96 individuals of Arabidopsis thaliana that included stock center accessions as well as a hierarchical sample from natural populations. Although A. thaliana is a selfing weed, the pattern of polymorphism in general agrees with what is expected for a widely distributed, sexually reproducing species. Linkage disequilibrium decays rapidly, within 50 kb. Variation is shared worldwide, although population structure and isolation by distance are evident. The data fail to fit standard neutral models in several ways. There is a genome-wide excess of rare alleles, at least partially due to selection. There is too much variation between genomic regions in the level of polymorphism. The local level of polymorphism is negatively correlated with gene density and positively correlated with segmental duplications. Because the data do not fit theoretical null distributions, attempts to infer natural selection from polymorphism data will require genome-wide surveys of polymorphism in order to identify anomalous regions. Despite this, our data support the utility of A. thaliana as a model for evolutionary functional genomics.

Arabidopsis↗

The signature of positive selection on standing genetic variation.

Considerable interest is focused on the use of polymorphism data to identify regions of the genome that underlie recent adaptations. These searches are guided by a simple model of positive selection, in which a mutation is favored as soon as it arises. This assumption may not be realistic, as environmental changes and range expansions may lead previously neutral or deleterious alleles to become beneficial. We examine what effect this mode of selection has on patterns of variation at linked neutral sites by implementing a new coalescent model of positive directional selection on standing variation. In this model, a neutral allele arises and drifts in the population, then at frequency f becomes beneficial, and eventually reaches fixation. Depending on the value of f, this scenario can lead to a large variance in allele frequency spectra and in levels of linkage disequilibrium at linked, neutral sites. In particular, for intermediate f, the beneficial substitution often leads to a loss of rare alleles--a pattern that differs markedly from the signature of directional selection currently relied on by researchers. These findings highlight the importance of an accurate characterization of the effects of positive selection, if we are to reliably identify recent adaptations from polymorphism data.

Biological Evolution↗

Estimating recombination rates using three-site likelihoods.

We introduce a new method for jointly estimating crossing-over and gene conversion rates using sequence polymorphism data. The method calculates probabilities for subsets of the data consisting of three segregating sites and then forms a composite likelihood by multiplying together the probabilities of many subsets. Simulations show that this new method performs better than previously proposed methods for estimating gene conversion rates, but that all methods require large amounts of data to provide reliable estimates. While existing methods can easily estimate an "average" gene conversion rate over many loci, they cannot reliably estimate gene conversion rates for a single region of the genome.

Computer Simulation↗

Comparative linkage-disequilibrium analysis of the beta-globin hotspot in primates.

Recombination rates vary both across the genome and between different species, but little information is available about the temporal and physical scales over which such rates change. To shed light on these questions, we performed a high-resolution analysis of a genomic region within the beta-globin gene cluster that is known to experience elevated recombination rates in humans. For this purpose, we developed new linkage disequilibrium-based methods that thoroughly search for subsets of the data with unusually high or unusually low estimated values of the population-recombination parameter (4Nr, where N is the effective population size and r is the crossover rate between adjacent base pairs). By resequencing a 15-kb segment in a human population sample, we were able to narrow the recombinational hotspot to a segment <2 kb in length that coincides with the beta-globin replication origin. In addition, we analyzed the orthologous region in samples of rhesus macaques and common chimpanzees. Whereas the analysis of the chimpanzee data is complicated by the sample structure, the macaque data imply that this region may not be a hotspot in that species. These results suggest a time scale for the evolution of hotspots in primates. Furthermore, they allow us to propose diverged sequence elements that may contribute to the differences in the recombinational landscape in the two species.

Animals↗

Assessing the performance of the haplotype block model of linkage disequilibrium.

Several recent studies have suggested that linkage disequilibrium (LD) in the human genome has a fundamentally "blocklike" structure. However, thus far there has been little formal assessment of how well the haplotype block model captures the underlying structure of LD. Here we propose quantitative criteria for assessing how blocklike LD is and apply these criteria to both real and simulated data. Analyses of several large data sets indicate that real data show a partial fit to the haplotype block model; some regions conform quite well, whereas others do not. Some improvement could be obtained by genotyping higher marker densities but not by increasing the number of samples. Nonetheless, although the real data are only moderately blocklike, our simulations indicate that, under a model of uniform recombination, the structure of LD would actually fit the block model much less well. Simulations of a model in which much of the recombination occurs in narrow hotspots provide a much better fit to the observed patterns of LD, suggesting that there is extensive fine-scale variation in recombination rates across the human genome.

Computer Simulation↗

Haplotype blocks and linkage disequilibrium in the human genome.

There is great interest in the patterns and extent of linkage disequilibrium (LD) in humans and other species. Characterizing LD is of central importance for gene-mapping studies and can provide insights into the biology of recombination and human demographic history. Here, we review recent developments in this field, including the recently proposed 'haplotype-block' model of LD. We describe some of the recent data in detail and compare the observed patterns to those seen in simulations.

Chromosome Mapping↗

Estimating ancestral population sizes and divergence times.

This article presents a new method for jointly estimating species divergence times and ancestral population sizes. The method improves on previous ones by explicitly incorporating intragenic recombination, by utilizing orthologous sequence data from closely related species, and by using a maximum-likelihood framework. The latter allows for efficient use of the available information and provides a way of assessing how much confidence we should place in the estimates. I apply the method to recently collected intergenic sequence data from humans and the great apes. The results suggest that the human-chimpanzee ancestral population size was four to seven times larger than the current human effective population size and that the current human effective population size is slightly >10,000. These estimates are similar to previous ones, and they appear relatively insensitive to assumptions about the recombination rates or mutation rates across loci.

Animals↗

Linkage disequilibrium patterns across a recombination gradient in African Drosophila melanogaster.

Previous multilocus surveys of nucleotide polymorphism have documented a genome-wide excess of intralocus linkage disequilibrium (LD) in Drosophila melanogaster and D. simulans relative to expectations based on estimated mutation and recombination rates and observed levels of diversity. These studies examined patterns of variation from predominantly non-African populations that are thought to have recently expanded their ranges from central Africa. Here, we analyze polymorphism data from a Zimbabwean population of D. melanogaster, which is likely to be closer to the standard population model assumptions of a large population with constant size. Unlike previous studies, we find that levels of LD are roughly compatible with expectations based on estimated rates of crossing over. Further, a detailed examination of genes in different recombination environments suggests that markers near the telomere of the X chromosome show considerably less linkage disequilibrium than predicted by rates of crossing over, suggesting appreciable levels of exchange due to gene conversion. Assuming that these populations are near mutation-drift equilibrium, our results are most consistent with a model that posits heterogeneity in levels of exchange due to gene conversion across the X chromosome, with gene conversion being a minor determinant of LD levels in regions of high crossing over. Alternatively, if levels of exchange due to gene conversion are not negligible in regions of high crossing over, our results suggest a marked departure from mutation-drift equilibrium (i.e., toward an excess of LD) in this Zimbabwean population. Our results also have implications for the dynamics of weakly selected mutations in regions of reduced crossing over.

Animals↗

Testing models of selection and demography in Drosophila simulans.

We analyze patterns of nucleotide variability at 15 X-linked loci and 14 autosomal loci from a North American population of Drosophila simulans. We show that there is significantly more linkage disequilibrium on the X chromosome than on chromosome arm 3R and much more linkage disequilibrium on both chromosomes than expected from estimates of recombination rates, mutation rates, and levels of diversity. To explore what types of evolutionary models might explain this observation, we examine a model of recurrent, nonoverlapping selective sweeps and a model of a recent drastic bottleneck (e.g., founder event) in the demographic history of North American populations of D. simulans. The simple sweep model is not consistent with the observed patterns of linkage disequilibrium nor with the observed frequencies of segregating mutations. Under a restricted range of parameter values, a simple bottleneck model is consistent with multiple facets of the data. While our results do not exclude some influence of selection on X vs. autosome variability levels, they suggest that demography alone may account for patterns of linkage disequilibrium and the frequency spectrum of segregating mutations in this population of D. simulans.

Animals↗