PubMed Health⌕ Search

Biomedical subjects

K Roeder

Publications and source records attributed to K Roeder.

At least 19 recordsLinked to original sources

Transmission/disequilibrium test meets measured haplotype analysis: family-based association analysis guided by evolution of haplotypes.

Family data teamed with the transmission/disequilibrium test (TDT), which simultaneously evaluates linkage and association, is a powerful means of detecting disease-liability alleles. To increase the information provided by the test, various researchers have proposed TDT-based methods for haplotype transmission. Haplotypes indeed produce more-definitive transmissions than do the alleles comprising them, and this tends to increase power. However, the larger number of haplotypes, relative to alleles at individual loci, tends to decrease power, because of the additional degrees of freedom required for the test. An optimal strategy would focus the test on particular haplotypes or groups of haplotypes. In this report we develop such an approach by combining the theory of TDT with that of measured haplotype analysis (MHA). MHA uses the evolutionary relationships among haplotypes to produce a limited set of hypothesis tests and to increase the interpretability of these tests. The theory of our approach, called the "evolutionary tree" (ET)-TDT, is developed for two cases: when haplotype transmission is certain and when it is not. Simulations show the ET-TDT can be more powerful than other proposed methods under reasonable conditions. More importantly, our results show that, when multiple polymorphisms are found within the gene, the ET-TDT can be useful for determining which polymorphisms affect liability.

Alleles↗

A Bayesian hierarchical model for allele frequencies.

Genetic epidemiological methodologies, such as linkage analysis, often require accurate estimates of allele frequencies. When studies involve multiple sub-populations with different evolutionary histories, accurate estimates can be difficult to obtain because the number of subjects per sub-population tends to be limited. Given allele counts for a collection of loci and sub-populations, we propose a Bayesian hierarchical model that extends existing empirical Bayesian approaches by allowing for explicit inclusion of prior information about both allele frequencies and inter-population divergence. We describe how such information can be derived from published data and then incorporated into the model via prior distributions for model parameters. By analysis of simulated data, we highlight how the hierarchical model, as implemented in the publicly available program AllDist, combines prior information with the observed data to refine allele frequency estimates.

Alleles↗

Unbiased methods for population-based association studies.

Large, population-based samples and large-scale genotyping are being used to evaluate disease/gene associations. A substantial drawback to such samples is the fact that population substructure can induce spurious associations between genes and disease. We review two methods, called genomic control (GC) and structured association (SA), that obviate many of the concerns about population substructure by using the features of the genomes present in the sample to correct for stratification. The GC approach exploits the fact that population substructure generates "over dispersion" of statistics used to assess association. By testing multiple polymorphisms throughout the genome, only some of which are pertinent to the disease of interest, the degree of overdispersion generated by population substructure can be estimated and taken into account. The SA approach assumes that the sampled population, although heterogeneous, is composed of subpopulations that are themselves homogeneous. By using multiple polymorphisms throughout the genome, this "latent class method" estimates the probability sampled individuals derive from each of these latent subpopulations. GC has the advantage of robustness, simplicity, and wide applicability, even to experimental designs such as DNA pooling. SA is a bit more complicated but has the advantage of greater power in some realistic settings, such as admixed populations or when association varies widely across subpopulations. It, too, is widely applicable. Both also have weaknesses, as elaborated in our review.

Analysis of Variance↗

Genomic control, a new approach to genetic-based association studies.

During the past decade, mutations affecting liability to human disease have been discovered at a phenomenal rate, and that rate is increasing. For the most part, however, those diseases have a relatively simple genetic basis. For diseases with a complex genetic and environmental basis, new approaches are needed to pave the way for more rapid discovery of genes affecting liability. One such approach exploits large, population-based samples and large-scale genotyping to evaluate disease/gene associations. A substantial drawback to such samples is the fact that population heterogeneity can induce spurious associations between genes and disease. We describe a method called genomic control (GC), which obviates many of the concerns about population substructure by using the features of the genomes present in the sample to correct for stratification. Two such approaches are now available. The GC approach exploits the fact that population substructure generate "overdispersion" of statistics used to assess association. By testing multiple polymorphisms throughout the genome, only some of which are pertinent to the disease of interest, the degree of overdispersion generated by population substructure can be estimated and taken into account. The other approach, called Structured Association (SA), assumes that the sampled population, while heterogeneous, is composed of subpopulations that are themselves homogeneous. By using multiple polymorphisms throughout the genome, SA probabilistically assigns sampled individuals to these latent subpopulations. We review in detail the overdispersion GC. In addition to outlining the published ideas on this method, we describe several extensions: quantitative trait studies and case-control studies with haplotypes and multiallelic markers. For each study design our goal is to achieve control similar to that obtained for a family-based study, but with the convenience found in a population-based design.

Alleles↗

Genome-wide distribution of linkage disequilibrium in the population of Palau and its implications for gene flow in Remote Oceania.

Linkage disequilibrium (LD) between alleles on the same human chromosome results from various evolutionary processes and is thus telling about the history of populations. Recently, LD has garnered substantial interest for its value to map and fine-map disease genes. We examine the distribution of LD between short tandem repeat alleles on autosomes and sex chromosomes in the Remote Oceanic population of Palau to evaluate whether the data are consistent with a recent hypothesis about the origins of genetic variation in Palau, specifically that the population experienced extensive male-biased gene flow following initial settlement. Consistent with evolutionary theory based on effective population size, LD between X-linked alleles is stochastically greater than LD between autosomal alleles, however, small but detectable LD occurs for autosomal markers separated by substantial distances. By contrast, while Y-linked alleles experience only one-third the effective population size of X-linked alleles, their mean value for pairwise LD is only slightly larger than X-linked alleles. For a small population known to experience at least two extreme bottlenecks, 56 six-locus Y haplotypes exhibit remarkable diversity (0.96), comparable to Y diversity of Europeans, however, autosomal and X-linked markers display significantly less diversity, as measured by heterozygosity (4.1% less). Palauan Y haplotypes also fall into distinct clusters, again unlike that of Europe. We argue these data are consistent with waves of male-biased gene flow.

Alleles↗

The power of genomic control.

Although association analysis is a useful tool for uncovering the genetic underpinnings of complex traits, its utility is diminished by population substructure, which can produce spurious association between phenotype and genotype within population-based samples. Because family-based designs are robust against substructure, they have risen to the fore of association analysis. Yet, if population substructure could be ignored, this robustness can come at the price of power. Unfortunately it is rarely evident when population substructure can be ignored. Devlin and Roeder recently have proposed a method, termed "genomic control" (GC), which has the robustness of family-based designs even though it uses population-based data. GC uses the genome itself to determine appropriate corrections for population-based association tests. Using the GC method, we contrast the power of two study designs, family trios (i.e., father, mother, and affected progeny) versus case-control. For analysis of trios, we use the TDT test. When population substructure is absent, we find GC is always more powerful than TDT; furthermore, contrary to previous results, we show that as a disease becomes more prevalent the discrepancy in power becomes more extreme. When population substructure is present, however, the results are more complex: TDT is more powerful when population substructure is substantial, and GC is more powerful otherwise. We also explore general issues of power and implementation of GC within the case-control setting and find that, economically, GC is at least comparable to and often less expensive than family-based methods. Therefore, GC methods should prove a useful complement to family-based methods for the genetic analysis of complex traits.

Alleles↗

Haplotype fine mapping by evolutionary trees.

To refine the location of a disease gene within the bounds provided by linkage analysis, many scientists use the pattern of linkage disequilibrium between the disease allele and alleles at nearby markers. We describe a method that seeks to refine location by analysis of "disease" and "normal" haplotypes, thereby using multivariate information about linkage disequilibrium. Under the assumption that the disease mutation occurs in a specific gap between adjacent markers, the method first combines parsimony and likelihood to build an evolutionary tree of disease haplotypes, with each node (haplotype) separated, by a single mutational or recombinational step, from its parent. If required, latent nodes (unobserved haplotypes) are incorporated to complete the tree. Once the tree is built, its likelihood is computed from probabilities of mutation and recombination. When each gap between adjacent markers is evaluated in this fashion and these results are combined with prior information, they yield a posterior probability distribution to guide the search for the disease mutation. We show, by evolutionary simulations, that an implementation of these methods, called "FineMap," yields substantial refinement and excellent coverage for the true location of the disease mutation. Moreover, by analysis of hereditary hemochromatosis haplotypes, we show that FineMap can be robust to genetic heterogeneity.

Algorithms↗

Flexible parametric measurement error models.

Inferences in measurement error models can be sensitive to modeling assumptions. Specifically, if the model is incorrect, the estimates can be inconsistent. To reduce sensitivity to modeling assumptions and yet still retain the efficiency of parametric inference, we propose using flexible parametric models that can accommodate departures from standard parametric models. We use mixtures of normals for this purpose. We study two cases in detail: a linear errors-in-variables model and a change-point Berkson model.

Biometry↗

Genomic control for association studies.

A dense set of single nucleotide polymorphisms (SNP) covering the genome and an efficient method to assess SNP genotypes are expected to be available in the near future. An outstanding question is how to use these technologies efficiently to identify genes affecting liability to complex disorders. To achieve this goal, we propose a statistical method that has several optimal properties: It can be used with case control data and yet, like family-based designs, controls for population heterogeneity; it is insensitive to the usual violations of model assumptions, such as cases failing to be strictly independent; and, by using Bayesian outlier methods, it circumvents the need for Bonferroni correction for multiple tests, leading to better performance in many settings while still constraining risk for false positives. The performance of our genomic control method is quite good for plausible effects of liability genes, which bodes well for future genetic analyses of complex disorders.

Bayes Theorem↗

The heritability of IQ.

IQ heritability, the portion of a population's IQ variability attributable to the effects of genes, has been investigated for nearly a century, yet it remains controversial. Covariance between relatives may be due not only to genes, but also to shared environments, and most previous models have assumed different degrees of similarity induced by environments specific to twins, to non-twin siblings (henceforth siblings), and to parents and offspring. We now evaluate an alternative model that replaces these three environments by two maternal womb environments, one for twins and another for siblings, along with a common home environment. Meta-analysis of 212 previous studies shows that our 'maternal-effects' model fits the data better than the 'family-environments' model. Maternal effects, often assumed to be negligible, account for 20% of covariance between twins and 5% between siblings, and the effects of genes are correspondingly reduced, with two measures of heritability being less than 50%. The shared maternal environment may explain the striking correlation between the IQs of twins, especially those of adult twins that were reared apart. IQ heritability increases during early childhood, but whether it stabilizes thereafter remains unclear. A recent study of octogenarians, for instance, suggests that IQ heritability either remains constant through adolescence and adulthood, or continues to increase with age. Although the latter hypothesis has recently been endorsed, it gathers only modest statistical support in our analysis when compared to the maternal-effects hypothesis. Our analysis suggests that it will be important to understand the basis for these maternal effects if ways in which IQ might be increased are to be identified.

Adolescent↗

A statistical model for locating regulatory regions in genomic DNA.

In addition to genes, chromosomal DNA contains sequences that serve as signals for turning on and off gene expression. These signals are thought to be distributed as clusters in the regulatory regions of genes. We develop a Bayesian model that views locating regulatory regions in genomic DNA as a change-point problem, with the beginning of regulatory and non-regulatory regions corresponding to the change points. The model is based on a hidden Markov chain. The data consist of nucleotide positions of protein-binding elements in a genomic DNA sequence. These positions are identified using a reference catalogue containing elements that interact with transcription factors implicated in controlling the expression of protein-encoding genes. Among the protein-binding elements in a genomic DNA sequence, the statistical model automatically selects those that tend to predict regulatory regions. We test the model using viral sequences that include known regulatory regions and provide the results obtained for human genomic DNA corresponding to the beta globin locus on chromosome 11.

Adenoviridae↗

Binning clones by hybridization with complex probes: statistical refinement of an inner product mapping method.

Molecular methods that use long-range information to solve genomics problems (i.e., top-down strategies) efficiently have become increasingly prominent in the genomics literature. One such method, an implementation of inner product mapping (IPM), uses noisy, long-range radiation hybrid (RH)/YAC overlap data and relatively noise-free RH/STS overlap data to localize clones to specific chromosomal regions. Because the molecular data are rarely noise-free, statistical models tailored to the top-down molecular methods make the methods far more effective. We develop two statistical models for IPM (or any other top-down strategy of similar form), a parametric logit model and a nonparametric order-restricted model, and show how these models can be implemented within a hierarchical Bayes framework. Using these models, we refine the chromosome 11 map reported in M. Perlin et al. (1995, Genomics 28: 315-327). Our analyses improve the IPM map, both in terms of successful localization of clones and in terms of the confidence with which they are localized.

Chromosome Mapping↗

Disequilibrium mapping: composite likelihood for pairwise disequilibrium.

The pattern of linkage disequilibrium between a disease locus and a set of marker loci has been shown to be a useful tool for geneticists searching for disease genes. Several methods have been advanced to utilize the pairwise disequilibrium between the disease locus and each of a set of marker loci. However, none of the methods take into account the information from all pairs simultaneously while also modeling the variability in the disequilibrium values due to the evolutionary dynamics of the population. We propose a Composite Likelihood (CL) model that has these features when the physical distances between the marker loci are known or can be approximated. In this instance, and assuming that there is a single disease mutation, the CL model depends on only three parameters, the recombination fraction between the disease locus and an arbitrary marker locus, theta, the age of the mutation, and a variance parameter. When the CL is maximized over a grid of theta, it provides a graph that can direct the search for the disease locus. We also show how the CL model can be generalized to account for multiple disease mutations. Evolutionary simulations demonstrate the power of the analyses, as well as their potential weaknesses. Finally, we analyze the data from two mapped diseases, cystic fibrosis and diastrophic dysplasia, finding that the CL method performs well in both cases.

Computer Simulation↗

Comments on the statistical aspects of the NRC's report on DNA typing.

The goal of the NRC report on DNA typing was to answer a "crescendo of questions concerning DNA typing," many of them in the areas of population genetics and statistics. Unfortunately, few of these questions were answered adequately. In lieu of answering these questions, the panel proposed another conservative method of forensic inference, the "ceiling principle." Aside from its extreme conservativeness, this new method is difficult to justify because it is based on inadequate population genetics and statistical theory. Moreover, in its ultimate implementation, the panel's method will depend on a population genetics study whose rationale is questionable. In this article, we elaborate some of the general comments we made about the NRC report in a recent article [1]. Specifically we cover three topics. First we question the statistical basis for the ceiling principle, showing that the empirical results that motivated the method are likely to be misinterpreted and showing, by power calculations, that the effects of population substructure cannot be substantial. Second, we show that the study design to determine "ceiling" allele frequencies has several undesirable statistical properties. Finally, we discuss the estimation of handling errors from the statistical perspective, a subject treated inadequately by the report.

DNA Fingerprinting↗