PubMed Health⌕ Search

Biomedical subjects

Montgomery Slatkin

Publications and source records attributed to Montgomery Slatkin.

At least 19 recordsLinked to original sources

Inference of population genetic parameters in metagenomics: a clean look at messy data.

Metagenomic projects generate short, overlapping fragments of DNA sequence, each deriving from a different individual. We report a new method for inferring the scaled mutation rate, theta = 2Neu, and the scaled exponential growth rate, R = Ner, from the site-frequency spectrum of these data while accounting for sequencing error via Phred quality scores. After obtaining maximum likelihood parameter estimates for theta and R, we calculate empirical Bayes quality scores reflecting the posterior probability that each apparently polymorphic site is truly polymorphic; these scores can then be used for other applications such as SNP discovery. For realistic parameter ranges, analytic and simulation results show our estimates to be essentially unbiased with tight confidence intervals. In contrast, choosing an arbitrary quality score cutoff (e.g., trimming reads) and ignoring further quality information during inference yields biased estimates with greater variance. We illustrate the use of our technique on a new project analyzing activated sludge from a lab-scale bioreactor seeded by a wastewater treatment plant.

Bacteria↗

Non-equilibrium theory of the allele frequency spectrum.

A forward diffusion equation describing the evolution of the allele frequency spectrum is presented. The influx of mutations is accounted for by imposing a suitable boundary condition. For a Wright-Fisher diffusion with or without selection and varying population size, the boundary condition is lim(x downward arrow0)xf(x,t)=thetarho(t), where f(.,t) is the frequency spectrum of derived alleles at independent loci at time t and rho(t) is the relative population size at time t. When population size and selection intensity are independent of time, the forward equation is equivalent to the backwards diffusion usually used to derive the frequency spectrum, but this approach allows computation of the time dependence of the spectrum both before an equilibrium is attained and when population size and selection intensity vary with time. From the diffusion equation, a set of ordinary differential equations for the moments of f(.,t) is derived and the expected spectrum of a finite sample is expressed in terms of those moments. The use of the forward equation is illustrated by considering neutral and selected alleles in a highly simplified model of human history. For example, it is shown that approximately 30% of the expected total heterozygosity of neutral loci is attributable to mutations that arose since the onset of population growth in roughly the last 150,000 years.

Gene Frequency↗

Multiplex amplification of the mammoth mitochondrial genome and the evolution of Elephantidae.

In studying the genomes of extinct species, two principal limitations are typically the small quantities of endogenous ancient DNA and its degraded condition, even though products of up to 1,600 base pairs (bp) have been amplified in rare cases. Using small overlapping polymerase chain reaction products, longer stretches of sequences or even whole mitochondrial genomes can be reconstructed, but this approach is limited by the number of amplifications that can be performed from rare samples. Thus, even from well-studied Pleistocene species such as mammoths, ground sloths and cave bears, no DNA sequences of more than about 1,000 bp have been reconstructed. Here we report the complete mitochondrial genome sequence of the Pleistocene woolly mammoth Mammuthus primigenius. We used about 200 mg of bone and a new approach that allows the simultaneous retrieval of multiple sequences from small amounts of degraded DNA. Our phylogenetic analyses show that the mammoth was more closely related to the Asian than to the African elephant. However, the divergence of mammoth, African and Asian elephants occurred over a short time, corresponding to only about 7% of the total length of the phylogenetic tree for the three evolutionary lineages.

Africa↗

The concordance of gene trees and species trees at two linked loci.

The gene genealogies of two linked loci in three species are analyzed using a series of Markov chain models. We calculate the probability that the gene tree of one locus is concordant with the species tree, given that the gene tree of the other locus is concordant. We define a threshold value of the recombination rate, r*, to be the rate for which the difference between the conditional probability of concordance and its asymptotic value is reduced to 5% of the initial difference. We find that, although r* depends in a complicated way on the times of speciation and effective population sizes, it is always relatively small, <10/N4, where N4 is the effective size of the species represented by the internal branch of the species tree. Consequently, the concordance of gene trees of neutral loci with the species tree is expected to be on roughly the same length scale on the chromosome as the extent of significant linkage disequilibrium within species unless the effective size of contemporary populations is very different from the effective sizes of their ancestral populations. Both balancing selection and selective sweeps can result in much longer genomic regions having concordant gene trees.

Animals↗

The geographic spread of the CCR5 Delta32 HIV-resistance allele.

The Delta32 mutation at the CCR5 locus is a well-studied example of natural selection acting in humans. The mutation is found principally in Europe and western Asia, with higher frequencies generally in the north. Homozygous carriers of the Delta32 mutation are resistant to HIV-1 infection because the mutation prevents functional expression of the CCR5 chemokine receptor normally used by HIV-1 to enter CD4+ T cells. HIV has emerged only recently, but population genetic data strongly suggest Delta32 has been under intense selection for much of its evolutionary history. To understand how selection and dispersal have interacted during the history of the Delta32 allele, we implemented a spatially explicit model of the spread of Delta32. The model includes the effects of sampling, which we show can give rise to local peaks in observed allele frequencies. In addition, we show that with modest gradients in selection intensity, the origin of the Delta32 allele may be relatively far from the current areas of highest allele frequency. The geographic distribution of the Delta32 allele is consistent with previous reports of a strong selective advantage (>10%) for Delta32 carriers and of dispersal over relatively long distances (>100 km/generation). When selection is assumed to be uniform across Europe and western Asia, we find support for a northern European origin and long-range dispersal consistent with the Viking-mediated dispersal of Delta32 proposed by G. Lucotte and G. Mercier. However, when we allow for gradients in selection intensity, we estimate the origin to be outside of northern Europe and selection intensities to be strongest in the northwest. Our results describe the evolutionary history of the Delta32 allele and establish a general methodology for studying the geographic distribution of selected alleles.

Alleles↗

The beta -globin recombinational hotspot reduces the effects of strong selection around HbC, a recently arisen mutation providing resistance to malaria.

Recombination is expected to reduce the effect of selection on the extent of linkage disequilibrium (LD), but the impact that recombinational hotspots have on sites linked to selected mutations has not been investigated. We empirically determine chromosomal linkage phase for 5.2 kb spanning the beta -globin gene and hotspot. We estimate that the HbC mutation, which is positively selected because of malaria, originated <5,000 years ago and that selection coefficients are 0.04-0.09. Despite strong selection and the recent origin of the HbC allele, recombination (crossing-over or gene conversion) is observed within 1 kb 5' of the selected site on more than one-third of the HbC chromosomes sampled. The rapid decay in LD upstream of the HbC allele demonstrates the large effect the ss-globin hotspot has in mitigating the effects of positive selection on linked variation.

Globins↗

The extent of linkage disequilibrium caused by selection on G6PD in humans.

The gene coding for glucose-6-phosphate dehydrogenase (G6PD) is subject to positive selection by malaria in some human populations. The G6PD A- allele, which is common in sub-Saharan Africa, is associated with deficient enzyme activity and protection from severe malaria. To delimit the impact of selection on patterns of linkage disequilibrium (LD) and nucleotide diversity, we resequenced 5.1 kb at G6PD and approximately 2-3 kb at each of eight loci in a 2.5-Mb region roughly centered on G6PD in a diverse sub-Saharan African panel of 51 unrelated men (including 20 G6PD A-, 11 G6PD A+, and 20 G6PD B chromosomes). The signature of selection is evident in the absence of genetic variation at G6PD and at three neighboring loci within 0.9 Mb from G6PD among all individuals bearing G6PD A- alleles. A genomic region of approximately 1.6 Mb around G6PD was characterized by long-range LD associated with the A- alleles. These patterns of nucleotide variability and LD suggest that G6PD A- is younger than previous age estimates and has increased in frequency in sub-Saharan Africa due to strong selection (0.1 < s < 0.2). These results also show that selection can lead to nonrandom associations among SNPs over great physical and genetic distances, even in African populations.

Alleles↗

Seeing ghosts: the effect of unsampled populations on migration rates estimated for sampled populations.

In 2004, the term 'ghost population' was introduced to summarize the effect of unsampled subpopulations that exchange migrants with other subpopulations that have been sampled. Estimated long-term migration rates among populations sampled will be affected by ghost populations. Although it would be convenient to be able to define an apparent migration matrix among sampled populations that incorporate the exchange of migrants with ghost populations, no such matrix can be defined in a way that predicts all features of the coalescent process for the true migration matrix. This paper shows that if the underlying migration matrix is symmetric, it is possible to define an apparent migration matrix among sampled subpopulations that predicts the same within-population and between-population homozygosities among sampled populations as is predicted by the true migration matrix. Application of this method shows that there is no simple relationship between true and apparent migration rates, nor is there a way to place an upper bound on the effect of ghost populations. In general, ghost populations can create the appearance of migration between subpopulations that do not actually exchange migrants. Comparison with published results from the application of the program, MIGRATE, shows that the apparent migration rates inferred with that program in a three-subpopulation model differ from those based on pairwise homozygosities. The apparent migration matrix determined by the method described in this paper probably represents the upper bound on the effect of ghost populations.

Homozygote↗

Breed distribution and history of canine mdr1-1Delta, a pharmacogenetic mutation that marks the emergence of breeds from the collie lineage.

A mutation in the canine multidrug resistance gene, MDR1, has previously been associated with drug sensitivities in two breeds from the collie lineage. We exploited breed phylogeny and reports of drug sensitivity to survey other purebred populations that might be genetically at risk. We found that the same allele, mdr1-1Delta, segregated in seven additional breeds, including two sighthounds that were not expected to share collie ancestry. A mutant haplotype that was conserved among affected breeds indicated that the allele was identical by descent. Based on breed histories and the extent of linkage disequilibrium, we conclude that all dogs carrying mdr1-1Delta are descendants of a dog that lived in Great Britain before the genetic isolation of breeds by registry (ca. 1873). The breed distribution and frequency of mdr1-1Delta have applications in veterinary medicine and selective breeding, whereas the allele's history recounts the emergence of formally recognized breeds from an admixed population of working sheepdogs.

Alleles↗

A population-genetic test of founder effects and implications for Ashkenazi Jewish diseases.

A founder effect can account for the presence of an allele at an unusually high frequency in an isolated population if the allele is selectively neutral and if all copies are identical by descent with a copy that either was carried by a founder individual or arose by mutation later. Here, a statistical test of both aspects of the founder-effect hypothesis is developed. The test is performed by a modified version of a program that implements the Slatkin-Bertorelle test of neutrality. The test is applied to several disease-associated alleles found predominantly in Ashkenazi Jews. Despite considerable uncertainty about the demographic history of Ashkenazi Jews and their ancestors, available genetic data are consistent with a founder effect resulting from a severe bottleneck in population size between a.d. 1100 and a.d. 1400 and an earlier bottleneck in a.d. 75, at the beginning of the Jewish Diaspora. The relatively high frequency of alleles causing four different lysosomal storage disorders, including Tay-Sachs disease and Gaucher disease, can be accounted for if the disease-associated alleles are recessive in their effects on reproductive fitness.

Adenomatous Polyposis Coli↗

Intense selection in an age-structured population.

In a population with overlapping generations, intense selection can perturb the age distribution and thus affect the rate of increase of an advantageous allele. We found that the age-specific nature of intense selection, such as that generated by many diseases, can affect the outcome of selection on loci, such as those conferring disease resistance. We also found that the temporal dynamics of selection alter the speed of evolution, particularly when selection is intense, and even more so when it is age-specific. We relate our model and results to selection for disease resistance, although the results have broader implications for inferences about past selection pressures in general.

Age Factors↗

Population-genetic basis of haplotype blocks in the 5q31 region.

We investigated patterns of nucleotide variation in the 5q31 region identified by Daly et al. as containing haplotype blocks, to determine whether the blocklike pattern requires the assumption of hotspots in recombination. Using extensive simulations that generate data matched to the Daly et al. data set in (a) the method of ascertainment of single-nucleotide polymorphisms, (b) the heterozygosity of ascertained markers, (c) the number of block boundaries, and (d) the diversity of haplotypes within blocks, we show that the patterns found in the Daly et al. data are not consistent with the assumption of uniform recombination in a population of constant size but are consistent either with the presence of hotspots in a population of constant size or with the absence of hotspots if there was a period of rapid population growth. We further show that estimates of local recombination rate can distinguish between population growth and hotspots as the primary cause of a blocklike pattern. Estimates of local recombination rates for the Daly et al. data do not indicate the presence of recombination hotspots.

Chromosome Mapping↗

Evaluating plague and smallpox as historical selective pressures for the CCR5-Delta 32 HIV-resistance allele.

The high frequency, recent origin, and geographic distribution of the CCR5-Delta 32 deletion allele together indicate that it has been intensely selected in Europe. Although the allele confers resistance against HIV-1, HIV has not existed in the human population long enough to account for this selective pressure. The prevailing hypothesis is that the selective rise of CCR5-Delta 32 to its current frequency can be attributed to bubonic plague. By using a population genetic framework that takes into account the temporal pattern and age-dependent nature of specific diseases, we find that smallpox is more consistent with this historical role.

Adolescent↗

On selecting markers for association studies: patterns of linkage disequilibrium between two and three diallelic loci.

Association studies depend on linkage disequilibrium (LD) between a causative mutation and linked marker loci. Selecting markers that give the best chance of showing useful levels of LD with the causative mutation will increase the chances of successfully detecting an association. This report examines the variation in the extent of LD between a disease locus and one or two diallelic marker loci (termed single nucleotide polymorphisms or SNPs). We use a simulation method based on the neutral coalescent in a population of variable size to find the distribution of LD as a function of allele frequencies, the recombination rate, and the population history. Given that LD exists, the allele frequencies determine if a site will be useful for detecting an association with the disease mutation. We show that there is extensive variation in LD even for closely linked loci, implying that several markers may be needed to detect a disease locus. The distribution of LD between common variants is strongly influenced by ancestral population size. We show that in general, best results will be obtained if the frequencies of marker alleles are at least as large as the frequency of the causative mutation. Haplotypes of two or more SNPs generally have a higher probability than individual SNPs of showing useful LD with a disease mutation, although exceptions are described.

Alleles↗

The loss of statistical power to distinguish populations when certain samples are ambiguous.

Case-control studies are used to map loci associated with a genetic disease. The usual case-control study tests for significant differences in frequencies of alleles at marker loci. In this paper, we consider the problem of comparing two or more marker loci simultaneously and testing for significant differences in haplotype rather than allele frequencies. We consider two situations. In the first, genotypes at marker loci are resolved into haplotypes by making use of biochemical methods or by genotyping family members. In the second, genotypes at marker loci are not resolved into haplotypes, but, by assuming random mating, haplotypes can be inferred using a likelihood method such as the expectation-maximization (EM) algorithm. We assume that a causative locus has two alleles with a multiplicative effect on the penetrance of a disease, with one allele increasing the penetrance by a factor pi. We find, for small values of pi-1 and large sample sizes, asymptotic results that predict the statistical power of a test for significant differences in haplotype frequencies between cases and a random sample of the population, both when haplotypes can be resolved and when haplotypes have to be inferred. The increase in power when haplotypes can be resolved can be expressed as a ratio R, which is the increase in sample size needed to achieve the same power when haplotypes are resolved over when they are not resolved. In general, R depends on the pattern of linkage disequilibrium between the causative allele and the marker haplotypes but is independent of the frequency of the causative allele and, to a first approximation, is independent of pi. For the special situation of two di-allelic marker loci, we obtain a simple expression for R and its upper bound.

Alleles↗

Multiplex relative risk and estimation of the number of loci underlying an inherited disease.

Knowledge of the number of causative loci is necessary to estimate the power of mapping studies of complex diseases. In the present article, we reexamine a theory developed by Risch and its implications for estimating the number L of causative loci affecting a complex inherited disease. We first show that methods based on Risch's analysis can produce estimates of L that are inconsistent with the observed population prevalence of the disease. We demonstrate this point by showing that the maximum-likelihood estimate for L produced by the method of Farrall and Holder for cleft lip/cleft palate data is not consistent with the prevalence under the multiplicative model. We show how to incorporate disease prevalence and develop a maximum-likelihood method for estimating L that uses the entire distribution of numbers of affected individuals in families containing an affected individual. This method avoids the potential inconsistencies of the Risch method and has greater precision. We apply our method to data on cleft lip/cleft palate and schizophrenia.

Bias↗

Likelihood-based disequilibrium mapping for two-marker haplotype data.

We report a theory that gives the sampling distribution of two-marker haplotypes that are linked to a rare disease mutation. The sampling distribution is generated with successive Monte Carlo realizations of the coalescence of the disease mutation having recombination and marker mutation events placed along the lineage. Given a sample of mutation-bearing, two-marker haplotypes, the maximum likelihood estimate of the location of the disease mutation can be calculated from the generated sampling distribution, provided that one knows enough about the population history in order to model it. The two-marker likelihood method is compared to a single-marker likelihood and a composite likelihood. The two-marker maximum likelihood gives smaller confidence intervals for the location of the disease locus than a comparable single-marker maximum likelihood. The composite likelihood can give biased results and the bias increases as the extent of linkage disequilibrium on mutation-bearing chromosomes decreases. Haplotype configurations exist for which the composite likelihood will fail to place the disease locus in the correct marker interval.

Genetic Diseases, Inborn↗