PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genotype Data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Direct analysis of unphased SNP genotype data in population-based association studies via Bayesian partition modelling of haplotypes.

We describe a novel method for assessing the strength of disease association with single nucleotide polymorphisms (SNPs) in a candidate gene or small candidate region, and for estimating the corresponding haplotype relative risks of disease, using unphased genotype data directly. We begin by estimating the relative frequencies of haplotypes consistent with observed SNP genotypes. Under the Bayesian partition model, we specify cluster centres from this set of consistent SNP haplotypes. The remaining haplotypes are then assigned to the cluster with the "nearest" centre, where distance is defined in terms of SNP allele matches. Within a logistic regression modelling framework, each haplotype within a cluster is assigned the same disease risk, reducing the number of parameters required. Uncertainty in phase assignment is addressed by considering all possible haplotype configurations consistent with each unphased genotype, weighted in the logistic regression likelihood by their probabilities, calculated according to the estimated relative haplotype frequencies. We develop a Markov chain Monte Carlo algorithm to sample over the space of haplotype clusters and corresponding disease risks, allowing for covariates that might include environmental risk factors or polygenic effects. Application of the algorithm to SNP genotype data in an 890-kb region flanking the CYP2D6 gene illustrates that we can identify clusters of haplotypes with similar risk of poor drug metaboliser (PDM) phenotype, and can distinguish PDM cases carrying different high-risk variants. Further, the results of a detailed simulation study suggest that we can identify positive evidence of association for moderate relative disease risks with a sample of 1,000 cases and 1,000 controls.

Algorithms↗

Search for haplotype interactions that influence susceptibility to type 1 diabetes, through use of unphased genotype data.

Type 1 diabetes is a T-cell-mediated chronic disease characterized by the autoimmune destruction of pancreatic insulin-producing beta cells and complete insulin deficiency. It is the result of a complex interrelation of genetic and environmental factors, most of which have yet to be identified. Simultaneous identification of these genetic factors, through use of unphased genotype data, has received increasing attention in the past few years. Several approaches have been described, such as the modified transmission/disequilibrium test procedure, the conditional extended transmission/disequilibrium test, and the stepwise logistic-regression procedure. These approaches are limited either by being restricted to family data or by ignoring so-called "haplotype interactions" between alleles. To overcome this limit, the present study provides a general method to identify, on the basis of unphased genotype data, the haplotype blocks that interact to define the risk for a complex disease. The principle underpinning the proposal is minimal entropy. The performance of our procedure is illustrated for both simulated and real data. In particular, for a set of Dutch type 1 diabetes data, our procedure suggests some novel evidence of the interactions between and within haplotype blocks that are across chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 11, 12, 15, 16, 17, 19, and 21. The results demonstrate that, by considering interactions between potential disease haplotype blocks, we may succeed in identifying disease-predisposing genetic variants that might otherwise have remained undetected.

Algorithms↗

SNPP: automating large-scale SNP genotype data management.

UNLABELLED: To manage high-throughput single nucleotide polymorphism (SNP) genotyping data efficiently, we developed a dynamic general database management system-SNPP (SNP Processor). It provides several functions, including data importing with comparison, Mendelian inheritance check within pedigrees, data compiling and exporting. Furthermore, SNPP may generate files for repeat genotyping and transform them into files that can be executed by a liquid handling system. AVAILABILITY: http://orclinux.creighton.edu/snpp/ CONTACT: lanjuanzhao@creighton.edu

Algorithms↗

Estimation of inbreeding coefficients from genotypic data on multiple alleles, and application to estimation of clonality in malaria parasites.

Methods for estimating probability of identity by descent (f) are derived for data on numbers of genotypes at single loci and at pairs of loci with many alleles at each locus. The methods are general, but are specifically applied to data on genotype frequencies in zygotes of the malaria parasite sampled from its mosquito host in order to find the extent of outcrossing in the parasite and the degree of clonality in populations. It is assumed that zygotes are the outcome either of gametes of the same clone, in which they are identical at all loci, or are products of two random, unrelated clones. From the estimate of f an effective number of clones per human host can also be derived. For Plasmodium falciparum from a Tanzanian village, estimates of f are 0.33 from data on zygote frequencies at two multiallelic loci, indicating that two-thirds of zygotes produce recombinant type.

Algorithms↗

Cluster analysis and association study of structured multilocus genotype data.

We propose an algorithm for testing association using structured multilocus genotype data. The algorithm implements the clustering of the data by a hierarchical clustering technique and a k-means algorithm. After clustering, the program analyzes all the clusters together using the Mantel-Haenszel (MH) test, by which common associations in the clusters are examined. To use the MH test, the number of subpopulations has to be determined. A method of cross-validation (CV) and the k-means algorithm are applied for estimating the number of subpopulations. The algorithm described was implemented in the computer program POPSTRUCT. In the simulation study, we found that when the two groups with different marker allele frequencies were combined, an inflation of the type I errors was observed. The inflation was more marked when the differences in the marker allele frequencies were larger, the difference in the minor allele frequencies at the disease locus was larger, and the genotype relative risk associated with the disease locus was higher. Our simulation study indicated that the MH test was efficient for decreasing type I errors and increasing the power compared with any test performed on each cluster. Then, we compared the results of STRUCTURE, a model-based method, and POPSTRUCT, a distance-based method. When two subgroups with different allele frequencies were mixed together at a high fixed ratio, POPSTRUCT was superior to STRUCTURE in classifying the combined population into the accurate clusters, each of which reflects one of the original groups.

Algorithms↗

Conversion of capillary electrophoresis microchip genotyping data for analysis with Genetic Profiler software.

The collection and conversion of 4-color fluorescent genotyping data from capillary array electrophoresis microchip devices and its conversion to a format easily and rapidly analyzed by Genetic Profiler genotyping software is presented. Microchip fluorescence intensity data are acquired and stored as 4-color tab-delimited text. These files are converted to electrophoretic signal data (ESD) files using a utility program (TEXT-to-ESD) written in C. TEXT-to-ESD generates an ESD file by converting text data to binary data and then appending a 632-byte ESD-file trailer. Up to 96 ESD files are then assembled into a run folder and imported into Genetic Profiler, where data are reduced to 4-color electropherograms and analyzed. In this manner, DNA fragment sizing data acquired with our high-speed electrophoretic microchip devices can be rapidly analyzed using robust commercial software. Additionally, the conversion program allows sizing of data with Genetic Profiler that have been preprocessed using other third-party software, such as BaseFinder.

Alleles↗

Genomic regions exhibiting positive selection identified from dense genotype data.

The allele frequency spectrum of polymorphisms in DNA sequences can be used to test for signatures of natural selection that depart from the expected frequency spectrum under the neutral theory. We observed a significant (P = 0.001) correlation between the Tajima's D test statistic in full resequencing data and Tajima's D in a dense, genome-wide data set of genotyped polymorphisms for a set of 179 genes. Based on this, we used a sliding window analysis of Tajima's D across the human genome to identify regions putatively subject to strong, recent, selective sweeps. This survey identified seven Contiguous Regions of Tajima's D Reduction (CRTRs) in an African-descent population (AD), 23 in a European-descent population (ED), and 29 in a Chinese-descent population (XD). Only four CRTRs overlapped between populations: three between ED and XD and one between AD and ED. Full resequencing of eight genes within six CRTRs demonstrated frequency spectra inconsistent with neutral expectations for at least one gene within each CRTR. Identification of the functional polymorphism (and/or haplotype) responsible for the selective sweeps within each CRTR may provide interesting insights into the strongest selective pressures experienced by the human genome over recent evolutionary history.

Black or African American↗

SNP Chart: an integrated platform for visualization and interpretation of microarray genotyping data.

UNLABELLED: SNP Chart is a Java application for the visualization and interpretation of microarray genotyping data primarily derived from arrayed primer extension-based chemistries. Spot intensity output files from microarray analysis tools are imported into SNP Chart, together with a multi-channel TIFF image of the original array experiment and a list of the actual single nucleotide polymorphisms (SNPs) being tested. Data from different and/or replicate probes that interrogate the same SNP, but that are scattered across the array grid, can be reassembled into a single chart format, specific for the SNP. This allows a quick and very effective 'visualization'/'quality control' of the data from multiple probes for the same SNP that can be easily interpreted and manually scored as a genotype. AVAILABILITY: http://www.snpchart.ca.

Computer Graphics↗

Impact of missing genotype data on Monte-Carlo simulation based haplotype analysis.

In the context of haplotype association analysis of unphased genotype data, methods based on Monte-Carlo simulations are often used to compensate for missing or inappropriate asymptotic theory. Moreover, such methods are an indispensable means to deal with multiple testing problems. We want to call attention to a potential trap in this usually useful approach: The simulation approach may lead to strongly inflated type I errors in the presence of different missing rates between cases and controls, depending on the chosen test statistic. Here, we consider four different testing strategies for haplotype analysis of case-control data. We recommend to interpret results for data sets with non-comparable distributions of missing genotypes with special caution, in case the test statistic is based on inferred haplotypes per individual. Moreover, our results are important for the conduction and interpretation of genome-wide association studies.

Case-Control Studies↗

A comparison of bayesian methods for haplotype reconstruction from population genotype data.

In this report, we compare and contrast three previously published Bayesian methods for inferring haplotypes from genotype data in a population sample. We review the methods, emphasizing the differences between them in terms of both the models ("priors") they use and the computational strategies they employ. We introduce a new algorithm that combines the modeling strategy of one method with the computational strategies of another. In comparisons using real and simulated data, this new algorithm outperforms all three existing methods. The new algorithm is included in the software package PHASE, version 2.0, available online (http://www.stat.washington.edu/stephens/software.html).

Algorithms↗

Comparison of haplotype inference methods using genotypic data from unrelated individuals.

OBJECTIVE: Haplotypes are gaining popularity in studies of human genetics because they contain more information than does a single gene locus. However, current high-throughput genotyping techniques cannot produce haplotype information. Several statistical methods have recently been proposed to infer haplotypes based on unphased genotypes at several loci. The accuracy, efficiency, and computational time of these methods have been under intense scrutiny. In this report, our aim was to evaluate haplotype inference methods for genotypic data from unrelated individuals. METHODS: We compared the performance of three haplotype inference methods that are currently in use--HAPLOTYPER, hap, and PHASE--by applying them to a large data set from unrelated individuals with known haplotypes. We also applied these methods to coalescent-based simulation studies using both constant size and exponential growth models. The performance of these methods, along with that of the expectation-maximization algorithm, was further compared in the context of an association study. RESULTS: While the algorithm implemented in the software PHASE was found to be the most accurate in both real and simulated data comparisons, all four methods produced good results in the association study.

Algorithms↗

[Gathering and evaluation of phenotype data of haemophilia A patients for correlation with genotype data].

Haemophilia A is caused by a genetic defect of the factor VIII gene resulting in complete or considerable functional loss of factor VIII molecule within blood. The high bleeding risk of patients can be prevented by intravenous injections of factor VIII protein. However, 25% of patients affected with severe haemophilia, develop factor VIII antibodies against the concentrate substituted. Within this study we try to comprise the phenotypic parameters (e. g. detailed documentation of disease course, basic laboratory values) and the therapy-associated data (e. g. applicated type and amount of factor VIII, number of substitutions, factor VIII recovery, inhibitor development and inhibitor elimination). We hope to identify differences of variable therapeutic treatments on course of disease as already identified for the factor VIII gene defects. At least we expect that certain mutations and mutation types, respectively, can be referred to typical phenotypes and similar course of treatment protocols.

Databases, Factual↗

Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies.

We describe extensions to the method of Pritchard et al. for inferring population structure from multilocus genotype data. Most importantly, we develop methods that allow for linkage between loci. The new model accounts for the correlations between linked loci that arise in admixed populations ("admixture linkage disequilibium"). This modification has several advantages, allowing (1) detection of admixture events farther back into the past, (2) inference of the population of origin of chromosomal regions, and (3) more accurate estimates of statistical uncertainty when linked loci are used. It is also of potential use for admixture mapping. In addition, we describe a new prior model for the allele frequencies within each population, which allows identification of subtle population subdivisions that were not detectable using the existing method. We present results applying the new methods to study admixture in African-Americans, recombination in Helicobacter pylori, and drift in populations of Drosophila melanogaster. The methods are implemented in a program, structure, version 2.0, which is available at http://pritch.bsd.uchicago.edu.

Algorithms↗

Estimating haplotype-disease associations with pooled genotype data.

The genetic dissection of complex human diseases requires large-scale association studies which explore the population associations between genetic variants and disease phenotypes. DNA pooling can substantially reduce the cost of genotyping assays in these studies, and thus enables one to examine a large number of genetic variants on a large number of subjects. The availability of pooled genotype data instead of individual data poses considerable challenges in the statistical inference, especially in the haplotype-based analysis because of increased phase uncertainty. Here we present a general likelihood-based approach to making inferences about haplotype-disease associations based on possibly pooled DNA data. We consider cohort and case-control studies of unrelated subjects, and allow arbitrary and unequal pool sizes. The phenotype can be discrete or continuous, univariate or multivariate. The effects of haplotypes on disease phenotypes are formulated through flexible regression models, which allow a variety of genetic hypotheses and gene-environment interactions. We construct appropriate likelihood functions for various designs and phenotypes, accommodating Hardy-Weinberg disequilibrium. The corresponding maximum likelihood estimators are approximately unbiased, normally distributed, and statistically efficient. We develop simple and efficient numerical algorithms for calculating the maximum likelihood estimators and their variances, and implement these algorithms in a freely available computer program. We assess the performance of the proposed methods through simulation studies, and provide an application to the Finland-United States Investigation of NIDDM Genetics Study. The results show that DNA pooling is highly efficient in studying haplotype-disease associations. As a by-product, this work provides valid and efficient methods for estimating haplotype-disease associations with unpooled DNA samples.

Algorithms↗

Cytochrome P450 1A1 polymorphism and childhood leukemia: an analysis of matched pairs case-control genotype data.

The association between the genotypic frequencies of the cytochrome P450 1A1 polymorphism and the risk of childhood leukemia is explored with the data from a matched case-control study. The data are displayed in a 3 x 3 case-control array, and the discordant pair counts are assessed for quasi-independence, homogeneity, and symmetry. This statistical approach is contrasted to the more typical analysis of matched data based on a conditional logistic model and estimated odds ratios. The statistical analysis of 175 matched pairs (part of a large study of potential environmental/genetic influences on the risk of childhood leukemia) showed no evidence of an association between cytochrome P450 1A1 genotype frequencies and case-control status.

Age Distribution↗

Little loss of information due to unknown phase for fine-scale linkage-disequilibrium mapping with single-nucleotide-polymorphism genotype data.

We present the results of a simulation study that indicate that true haplotypes at multiple, tightly linked loci often provide little extra information for linkage-disequilibrium fine mapping, compared with the information provided by corresponding genotypes, provided that an appropriate statistical analysis method is used. In contrast, a two-stage approach to analyzing genotype data, in which haplotypes are inferred and then analyzed as if they were true haplotypes, can lead to a substantial loss of information. The study uses our COLDMAP software for fine mapping, which implements a Markov chain-Monte Carlo algorithm that is based on the shattered coalescent model of genetic heterogeneity at a disease locus. We applied COLDMAP to 100 replicate data sets simulated under each of 18 disease models. Each data set consists of haplotype pairs (diplotypes) for 20 SNPs typed at equal 50-kb intervals in a 950-kb candidate region that includes a single disease locus located at random. The data sets were analyzed in three formats: (1). as true haplotypes; (2). as haplotypes inferred from genotypes using an expectation-maximization algorithm; and (3). as unphased genotypes. On average, true haplotypes gave a 6% gain in efficiency compared with the unphased genotypes, whereas inferring haplotypes from genotypes led to a 20% loss of efficiency, where efficiency is defined in terms of root mean integrated square error of the location of the disease locus. Furthermore, treating inferred haplotypes as if they were true haplotypes leads to considerable overconfidence in estimates, with nominal 50% credibility intervals achieving, on average, only 19% coverage. We conclude that (1). given appropriate statistical analyses, the costs of directly measuring haplotypes will rarely be justified by a gain in the efficiency of fine mapping and that (2). a two-stage approach of inferring haplotypes followed by a haplotype-based analysis can be very inefficient for fine mapping, compared with an analysis based directly on the genotypes.

Algorithms↗

Haplotype analysis in the presence of informatively missing genotype data.

It is common to have missing genotypes in practical genetic studies, but the exact underlying missing data mechanism is generally unknown to the investigators. Although some statistical methods can handle missing data, they usually assume that genotypes are missing at random, that is, at a given marker, different genotypes and different alleles are missing with the same probability. These include those methods on haplotype frequency estimation and haplotype association analysis. However, it is likely that this simple assumption does not hold in practice, yet few studies to date have examined the magnitude of the effects when this simplifying assumption is violated. In this study, we demonstrate that the violation of this assumption may lead to serious bias in haplotype frequency estimates, and haplotype association analysis based on this assumption can induce both false-positive and false-negative evidence of association. To address this limitation in the current methods, we propose a general missing data model to characterize missing data patterns across a set of two or more markers simultaneously. We prove that haplotype frequencies and missing data probabilities are identifiable if and only if there is linkage disequilibrium between these markers under our general missing data model. Simulation studies on the analysis of haplotypes consisting of two single nucleotide polymorphisms illustrate that our proposed model can reduce the bias both for haplotype frequency estimates and association analysis due to incorrect assumption on the missing data mechanism. Finally, we illustrate the utilities of our method through its application to a real data set.

Algorithms↗

Haplotype effects on human survival: logistic regression models applied to unphased genotype data.

Haplotype based linkage disequilibrium (LD) mapping exhibits higher power than the single locus approach because it makes use of the LD information contained in the flanking markers. New statistical methods have been proposed to help to infer haplotype effects on human diseases using multi-locus genotype data collected from unrelated individuals. In this paper, we introduce a statistical procedure for measuring haplotype effects on human survival using the popular logistic regression model with haplotype based parameterizations. By modeling haplotype frequency as a function of age, our model infers haplotype effects by estimating and testing the slope parameters under different genetic mechanisms (multiplicative, dominant, or recessive). In addition, by estimating the sex-specific slope parameters, our model allows the detection of sex-specific haplotype effects or haplotype-sex interactions. As an example, we apply our model to an empirical dataset on a stress related gene, interleukin-6, to look for haplotypes that affect individual survival and for haplotype-sex interactions. We show that our logistic regression based haplotype model can be a helpful tool for researchers interested in the genetics of human aging and longevity.

Female↗