PubMed Health⌕ Search

Biomedical subjects

Paul Fearnhead

Publications and source records attributed to Paul Fearnhead.

12 recordsLinked to original sources

SequenceLDhot: detecting recombination hotspots.

MOTIVATION: There is much local variation in recombination rates across the human genome--with the majority of recombination occurring in recombination hotspots--short regions of around approximately 2 kb in length that have much higher recombination rates than neighbouring regions. Knowledge of this local variation is important, e.g. in the design and analysis of association studies for disease genes. Population genetic data, such as that generated by the HapMap project, can be used to infer the location of these hotspots. We present a new, efficient and powerful method for detecting recombination hotspots from population data. RESULTS: We compare our method with four current methods for detecting hotspots. It is orders of magnitude quicker, and has greater power, than two related approaches. It appears to be more powerful than HotspotFisher, though less accurate at inferring the precise positions of the hotspot. It was also more powerful than LDhot in some situations: particularly for weaker hotspots (10-40 times the background rate) when SNP density is lower (< 1/kb). AVAILABILITY: Program, data sets, and full details of results are available at: http://www.maths.lancs.ac.uk/~fearnhea/Hotspot.

Algorithms↗

Perfect simulation from nonneutral population genetic models: variable population size and population subdivision.

We show how the idea of monotone coupling from the past can produce simple algorithms for simulating samples at a nonneutral locus under a range of demographic models. We specifically consider a biallelic locus and either a general variable population size mode or a general migration model for population subdivision. We investigate the effect of demography on the efficacy of selection and the effect of selection on genetic divergence between populations.

Algorithms↗

The stationary distribution of allele frequencies when selection acts at unlinked loci.

We consider population genetics models where selection acts at a set of unlinked loci. It is known that if the fitness of an individual is multiplicative across loci, then these loci are independent. We consider general selection models, but assume parent-independent mutation at each locus. For such a model, the joint stationary distribution of allele frequencies is proportional to the stationary distribution under neutrality multiplied by a known function of the mean fitness of the population. We further show how knowledge of this stationary distribution enables direct simulation of the genealogy of a sample at a single-locus. For a specific selection model appropriate for complex disease genes, we use simulation to determine what features of the genealogy differ between our general selection model and a multiplicative model.

Analysis of Variance↗

A novel method with improved power to detect recombination hotspots from polymorphism data reveals multiple hotspots in human genes.

We introduce a new method for detection of recombination hotspots from population genetic data. This method is based on (a) defining an (approximate) penalized likelihood for how recombination rate varies with physical position and (b) maximizing this penalized likelihood over possible sets of recombination hotspots. Simulation results suggest that this is a more powerful method for detection of hotspots than are existing methods. We apply the method to data from 89 genes sequenced in African American and European American populations. We find many genes with multiple hotspots, and some hotspots show evidence of being population-specific. Our results suggest that hotspots are randomly positioned within genes and could be as frequent as one per 30 kb.

Black People↗

Maximum-likelihood estimation of coalescence times in genealogical trees.

We develop a method for maximum-likelihood estimation of coalescence times in genealogical trees, based on population genetics data. For this purpose, a Viterbi-type algorithm is constructed to maximize the joint likelihood of the coalescence times. Marginal confidence intervals for the coalescence times based on the profile likelihoods are also computed. Our method of finding MLEs and calculating C.I.'s appears to be more accurate than alternative numerical maximization methods, and maximum-likelihood inference appears to be more accurate than other existing model-free approaches to estimating coalescent times. We demonstrate the method on two different data sets: human Y chromosome DNA data and fungus DNA data.

Algorithms↗

Analysis of recombination in Campylobacter jejuni from MLST population data.

We analyze recombination in C. jejuni using MLST data from isolates taken from wild birds, cattle, wild rabbits, and water in a 100-km2 study region in Cheshire, UK. We use a recent approximate likelihood method for inference, based on combining likelihood information from all pairs of segregating (polymorphic) sites in the data. We find substantial evidence for recombination, but only for recombination with short tract lengths, of around 225-750 bp. We estimate that the rate of recombination is of a similar magnitude to the rate of mutation.

Animals↗

A comparison of three estimators of the population-scaled recombination rate: accuracy and robustness.

We have performed simulations to assess the performance of three population genetics approximate-likelihood methods in estimating the population-scaled recombination rate from sequence data. We measured performance in two ways: accuracy when the sequence data were simulated according to the (simplistic) standard model underlying the methods and robustness to violations of many different aspects of the standard model. Although we found some differences between the methods, performance tended to be similar for all three methods. Despite the fact that the methods are not robust to violations of the underlying model, our simulations indicate that patterns of relative recombination rates should be inferred reasonably well even if the standard model does not hold. In addition, we assess various techniques for improving the performance of approximate-likelihood methods. In particular we find that the composite-likelihood method of Hudson (2001) can be improved by including log-likelihood contributions only for pairs of sites that are separated by some prespecified distance.

Computer Simulation↗

Spatial epidemiology and natural population structure of Campylobacter jejuni colonizing a farmland ecosystem.

Recent progress in determining the population structure of Campylobacter jejuni, and discerning associations between genotypes and specific niches, has emphasized the shortfall in our understanding of the ecology and epidemiology of this bacterium. We examined the natural structure of the C. jejuni community associated with cattle farmland in the UK by structured spatiotemporal sampling of habitats, including livestock and wild animal faeces, environmental water and soil, over a 10-week period within a 100 km2 area. A total of 172 isolates were characterized using multilocus sequence typing into 65 sequence types (STs). Isolates from cattle faeces were significantly over-represented in the ST-61 complex, whereas isolates from wildlife faeces and water were more likely to belong to the ST-45 complex and a number of unusual STs, many of which were first encountered during this study. Sampling within a narrow spatiotemporal window permitted the application of novel statistical methods exploring the relationship between the genetic relatedness and spatial separation of isolates. This approach showed that isolates from the same sampling squares and squares separated by <1.0 km were genetically more similar than isolates separated by greater distances. Our study demonstrates the potential of multilocus sequence typing combined with spatial modelling in exploring natural transmission pathways for C. jejuni.

Agriculture↗

Application of coalescent methods to reveal fine-scale rate variation and recombination hotspots.

There has been considerable recent interest in understanding the way in which recombination rates vary over small physical distances, and the extent of recombination hotspots, in various genomes. Here we adapt, apply, and assess the power of recently developed coalescent-based approaches to estimating recombination rates from sequence polymorphism data. We apply full-likelihood estimation to study rate variation in and around a well-characterized recombination hotspot in humans, in the beta-globin gene cluster, and show that it provides similar estimates, consistent with those from sperm studies, from two populations deliberately chosen to have different demographic and selectional histories. We also demonstrate how approximate-likelihood methods can be used to detect local recombination hotspots from genomic-scale SNP data. In a simulation study based on 80 100-kb regions, these methods detect 43 out of 60 hotspots (ranging from 1 to 2 kb in size), with only two false positives out of 2000 subregions that were tested for the presence of a hotspot. Our study suggests that new computational tools for sophisticated analysis of population diversity data are valuable for hotspot detection and fine-scale mapping of local recombination rates.

DNA↗

Ancestral processes for non-neutral models of complex diseases.

We consider non-neutral models for unlinked loci, where the fitness of a chromosome or individual is not multiplicative across loci. Such models are suitable for many complex diseases, where there are gene-interactions. We derive a genealogical process for such models, called the complex selection graph (CSG). This coalescent-type process is related to the ancestral selection graph, and is derived from the ancestral influence graph by considering the limit as the recombination rate between loci gets large. We analyse the CSG both theoretically and via simulation. The main results are that the gene-interactions do not produce linkage disequilibrium, but do produce dependencies in allele frequencies between loci. For small selection rates, the distributions of the genealogy and the allele frequencies at a single locus are well-approximated by their distributions under a single locus model, where the fitness of each allele is the average of the true fitnesses of that allele with respect to the distribution of alleles at other loci.

Genetic Predisposition to Disease↗

Consistency of estimators of the population-scaled recombination rate.

We consider (approximate) likelihood methods for estimating the population-scaled recombination rate from population genetic data. We show that the dependence between the data from two regions of a chromosome decays inversely with the amount of recombination between the two regions. We use this result to show that the maximum likelihood estimator (mle) for the recombination rate, based on the composite likelihood of Fearnhead and Donnelly, is consistent. We also consider inference based on the pairwise likelihood of Hudson. We consider two approximations to this likelihood, and prove that the mle based on one of these approximations is consistent, while the mle based on the other approximation (which is used by McVean, Awadalla and Fearnhead) is not.

Evolution, Molecular↗

A coalescent-based method for detecting and estimating recombination from gene sequences.

Determining the amount of recombination in the genealogical history of a sample of genes is important to both evolutionary biology and medical population genetics. However, recurrent mutation can produce patterns of genetic diversity similar to those generated by recombination and can bias estimates of the population recombination rate. Hudson 2001 has suggested an approximate-likelihood method based on coalescent theory to estimate the population recombination rate, 4N(e)r, under an infinite-sites model of sequence evolution. Here we extend the method to the estimation of the recombination rate in genomes, such as those of many viruses and bacteria, where the rate of recurrent mutation is high. In addition, we develop a powerful permutation-based method for detecting recombination that is both more powerful than other permutation-based methods and robust to misspecification of the model of sequence evolution. We apply the method to sequence data from viruses, bacteria, and human mitochondrial DNA. The extremely high level of recombination detected in both HIV1 and HIV2 sequences demonstrates that recombination cannot be ignored in the analysis of viral population genetic data.

Animals↗