PubMed Health⌕ Search

Biomedical subjects

Jonathan Marchini

Publications and source records attributed to Jonathan Marchini.

8 recordsLinked to original sources

A model-based approach to capture genetic variation for future association studies.

Genome-wide association studies are still constrained by the cost of genotyping. For this reason, the selection of a reduced set of markers or tags able to capture a significant proportion of the genetic variation is an important aspect of these studies. Most tagging SNP selection methods have been successful in capturing the genetic variation of the data from which the tags have been chosen. However, when these tags are used in an independent data set, a significant proportion of the remaining SNPs (non-tags) are not captured and, in most cases, there is no information on which SNPs are captured. We propose to use a probabilistic model to predict the non-tags based on a set of tags, as a way to capture genetic variation. An important advantage of this method is that it directly predicts the genotype of the non-tags with which we can test for association with the phenotype and which could help to elucidate the location of genes responsible for increasing disease susceptibility. Additionally, this method provides an estimate of the probabilities with which the predictions are made, which reflects the confidence of the probabilistic model. We also propose new methods to select the tagging SNPs. We empirically show by using HapMap data that our approach is able to capture significantly more genetic variation than methods based solely on a pairwise LD measure.

Algorithms↗

A high-resolution HLA and SNP haplotype map for disease association studies in the extended human MHC.

The proteins encoded by the classical HLA class I and class II genes in the major histocompatibility complex (MHC) are highly polymorphic and are essential in self versus non-self immune recognition. HLA variation is a crucial determinant of transplant rejection and susceptibility to a large number of infectious and autoimmune diseases. Yet identification of causal variants is problematic owing to linkage disequilibrium that extends across multiple HLA and non-HLA genes in the MHC. We therefore set out to characterize the linkage disequilibrium patterns between the highly polymorphic HLA genes and background variation by typing the classical HLA genes and >7,500 common SNPs and deletion-insertion polymorphisms across four population samples. The analysis provides informative tag SNPs that capture much of the common variation in the MHC region and that could be used in disease association studies, and it provides new insight into the evolutionary dynamics and ancestral origins of the HLA loci and their haplotypes.

Genetic Predisposition to Disease↗

Two-stage two-locus models in genome-wide association.

Studies in model organisms suggest that epistasis may play an important role in the etiology of complex diseases and traits in humans. With the era of large-scale genome-wide association studies fast approaching, it is important to quantify whether it will be possible to detect interacting loci using realistic sample sizes in humans and to what extent undetected epistasis will adversely affect power to detect association when single-locus approaches are employed. We therefore investigated the power to detect association for an extensive range of two-locus quantitative trait models that incorporated varying degrees of epistasis. We compared the power to detect association using a single-locus model that ignored interaction effects, a full two-locus model that allowed for interactions, and, most important, two two-stage strategies whereby a subset of loci initially identified using single-locus tests were analyzed using the full two-locus model. Despite the penalty introduced by multiple testing, fitting the full two-locus model performed better than single-locus tests for many of the situations considered, particularly when compared with attempts to detect both individual loci. Using a two-stage strategy reduced the computational burden associated with performing an exhaustive two-locus search across the genome but was not as powerful as the exhaustive search when loci interacted. Two-stage approaches also increased the risk of missing interacting loci that contributed little effect at the margins. Based on our extensive simulations, our results suggest that an exhaustive search involving all pairwise combinations of markers across the genome might provide a useful complement to single-locus scans in identifying interacting loci that contribute to moderate proportions of the phenotypic variance.

Alleles↗

A comparison of phasing algorithms for trios and unrelated individuals.

Knowledge of haplotype phase is valuable for many analysis methods in the study of disease, population, and evolutionary genetics. Considerable research effort has been devoted to the development of statistical and computational methods that infer haplotype phase from genotype data. Although a substantial number of such methods have been developed, they have focused principally on inference from unrelated individuals, and comparisons between methods have been rather limited. Here, we describe the extension of five leading algorithms for phase inference for handling father-mother-child trios. We performed a comprehensive assessment of the methods applied to both trios and to unrelated individuals, with a focus on genomic-scale problems, using both simulated data and data from the HapMap project. The most accurate algorithm was PHASE (v2.1). For this method, the percentages of genotypes whose phase was incorrectly inferred were 0.12%, 0.05%, and 0.16% for trios from simulated data, HapMap Centre d'Etude du Polymorphisme Humain (CEPH) trios, and HapMap Yoruban trios, respectively, and 5.2% and 5.9% for unrelated individuals in simulated data and the HapMap CEPH data, respectively. The other methods considered in this work had comparable but slightly worse error rates. The error rates for trios are similar to the levels of genotyping error and missing data expected. We thus conclude that all the methods considered will provide highly accurate estimates of haplotypes when applied to trio data sets. Running times differ substantially between methods. Although it is one of the slowest methods, PHASE (v2.1) was used to infer haplotypes for the 1 million-SNP HapMap data set. Finally, we evaluated methods of estimating the value of r(2) between a pair of SNPs and concluded that all methods estimated r(2) well when the estimated value was >or=0.8.

Algorithms↗

Genome-wide strategies for detecting multiple loci that influence complex diseases.

After nearly 10 years of intense academic and commercial research effort, large genome-wide association studies for common complex diseases are now imminent. Although these conditions involve a complex relationship between genotype and phenotype, including interactions between unlinked loci, the prevailing strategies for analysis of such studies focus on the locus-by-locus paradigm. Here we consider analytical methods that explicitly look for statistical interactions between loci. We show first that they are computationally feasible, even for studies of hundreds of thousands of loci, and second that even with a conservative correction for multiple testing, they can be more powerful than traditional analyses under a range of models for interlocus interactions. We also show that plausible variations across populations in allele frequencies among interacting loci can markedly affect the power to detect their marginal effects, which may account in part for the well-known difficulties in replicating association results. These results suggest that searching for interactions among genetic loci can be fruitfully incorporated into analysis strategies for genome-wide association studies.

Alleles↗

The effects of human population structure on large genetic association studies.

Large-scale association studies hold substantial promise for unraveling the genetic basis of common human diseases. A well-known problem with such studies is the presence of undetected population structure, which can lead to both false positive results and failures to detect genuine associations. Here we examine approximately 15,000 genome-wide single-nucleotide polymorphisms typed in three population groups to assess the consequences of population structure on the coming generation of association studies. The consequences of population structure on association outcomes increase markedly with sample size. For the size of study needed to detect typical genetic effects in common diseases, even the modest levels of population structure within population groups cannot safely be ignored. We also examine one method for correcting for population structure (Genomic Control). Although it often performs well, it may not correct for structure if too few loci are used and may overcorrect in other settings, leading to substantial loss of power. The results of our analysis can guide the design of large-scale association studies.

Genetic Markers↗

Comparing methods of analyzing fMRI statistical parametric maps.

Approaches for the analysis of statistical parametric maps (SPMs) can be crudely grouped into three main categories in which different philosophies are applied to delineate activated regions. These being type I error control thresholding, false discovery rate (FDR) control thresholding and posterior probability thresholding. To better understand the properties of these main approaches, we carried out a simulation study to compare the approaches as they would be used on real data sets. Using default settings, we find that posterior probability thresholding is the most powerful approach, and type I error control thresholding provides the lowest levels of type I error. False discovery rate control thresholding performs in between the other approaches for both these criteria, although for some parameter settings this approach can approximate the performance of posterior probability thresholding. Based on these results, we discuss the relative merits of the three approaches in an attempt to decide upon an optimal approach. We conclude that viewing the problem of delineating areas of activation as a classification problem provides a highly interpretable framework for comparing the methods. Within this framework, we highlight the role of the loss function, which explicitly penalizes the types of errors that may occur in a given analysis.

Bayes Theorem↗

Intra-individual variation in resting metabolic rate during the menstrual cycle.

Little information exists on the extent of day-to-day intra-individual variation in resting metabolic rate (RMR) in women. The present study has investigated the intra-individual variation in RMR of women during the menstrual cycle. Nineteen women (naturally cycling non-pill users) were recruited to the study. Anthropometric and RMR measurements were taken at least three times per week for the duration of one complete menstrual cycle; measurements were taken for a second, consecutive cycle in eight of the nineteen subjects. RMR was measured by indirect calorimetry using a ventilated hood system under standardized conditions. The measurements made throughout each complete menstrual cycle were averaged and the levels of inter- and intra-individual variation in RMR were assessed by determining the CV (%). Mean RMR of the group was 5686 (sd 674) kJ/d; inter-individual variation in RMR was 11.8 %. There were wide differences in the intra-individual variation in RMR of women (CV range 1.7-10.4 %). The CV in ten subjects was small (2-4 %), while the CV in nine women was high (5-10 %), indicating a significant variation in RMR during the menstrual cycle in certain subjects. Using statistical models, it has been shown that there was a significant effect on RMR due to a subject-specific level of variability; this was the case even when accounting for a possible training effect. In conclusion, the findings from our present study show that RMR cannot be assumed to be 'stable' in all women. The implications of intra-individual variation in RMR and its impact on energy balance needs further research.

Adult↗