PubMed Health⌕ Search

Biomedical subjects

Chad Haynes

Publications and source records attributed to Chad Haynes.

14 recordsLinked to original sources

The effects of SNP genotyping errors on the power of the Cochran-Armitage linear trend test for case/control association studies.

The questions addressed in this paper are: What single nucleotide polymorphism (SNP) genotyping errors are most costly, in terms of minimum sample size necessary (MSSN) to maintain constant asymptotic power and significance level, when performing case-control studies of genetic association applying the Cochran-Armitage trend test? And which trend test or chi2 test is more powerful under standard genetic models with genotyping errors? Our strategy is to expand the non-centrality parameter of the asymptotic distribution of the trend test to approximate the MSSN using a Taylor series linear in the genotyping error rates. We apply our strategy to example scenarios that assume recessive, dominant, additive, or over-dominant disease models. The most costly errors are recording the more common homozygote as the less common homozygote, and the more common homozygote as the heterozygote, with MSSN that become indefinitely large as the minor SNP allele frequency approaches zero. Misclassifying the heterozygote as the less common homozygote is costly when using the recessive trend test on data from a recessive model. The chi2 test has power close to, but less than, the optimal trend test and is never dominated over all genetic models studied by any specific trend test.

Case-Control Studies↗

Intrinsic disorder is a common feature of hub proteins from four eukaryotic interactomes.

Recent proteome-wide screening approaches have provided a wealth of information about interacting proteins in various organisms. To test for a potential association between protein connectivity and the amount of predicted structural disorder, the disorder propensities of proteins with various numbers of interacting partners from four eukaryotic organisms (Caenorhabditis elegans, Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens) were investigated. The results of PONDR VL-XT disorder analysis show that for all four studied organisms, hub proteins, defined here as those that interact with > or = 10 partners, are significantly more disordered than end proteins, defined here as those that interact with just one partner. The proportion of predicted disordered residues, the average disorder score, and the number of predicted disordered regions of various lengths were higher overall in hubs than in ends. A binary classification of hubs and ends into ordered and disordered subclasses using the consensus prediction method showed a significant enrichment of wholly disordered proteins and a significant depletion of wholly ordered proteins in hubs relative to ends in worm, fly, and human. The functional annotation of yeast hubs and ends using GO categories and the correlation of these annotations with disorder predictions demonstrate that proteins with regulation, transcription, and development annotations are enriched in disorder, whereas proteins with catalytic activity, transport, and membrane localization annotations are depleted in disorder. The results of this study demonstrate that intrinsic structural disorder is a distinctive and common characteristic of eukaryotic hub proteins, and that disorder may serve as a determinant of protein interactivity.

Amino Acids↗

Increase in linkage information by stratification of pedigree data into gold-standard and standard diagnoses: application to the NIMH Alzheimer Disease Genetics Initiative Dataset.

Patients diagnosed with a standard clinical method (subject to misclassification error) are often combined with patients diagnosed with a gold-standard method (with zero or very small misclassification error) in family-based studies of complex disease. For example, non-autopsied patients (NAP) are often included along with autopsy-proven (AP) patients in family-based studies of complex diseases, such as Alzheimer's disease (AD). Theoretical and simulation studies suggest that certain misclassification errors can result in severe reduction of power in genetic linkage and association analyses and that phenotype (or diagnostic) error can produce misleading results. Morton's test for heterogeneity can identify genomic regions where error may have led to loss in power. We applied this test to pedigree data from the NIMH Alzheimer's Disease Genetics Initiative Database separated into AP and NAP pedigrees. Morton's test identified one highly significant region of heterogeneity on chromosome 2. The source of the heterogeneity was due to significant indication of linkage in the AP pedigrees at position 109 cM (p value = 6.68 x 10(-5)) with no indication in the NAP pedigrees. Furthermore, Morton's test showed no evidence for heterogeneity on chromosome 19 in early-onset pedigrees that showed highly significant evidence for linkage in other published reports. These results suggest that supplementing linkage analysis with Morton's test can be usefully applied to genetic data sets that have AP and NAP samples, or other sample mixtures that include a 'gold standard' subgroup with reduced error rate, to increase power to detect linkage in the presence of diagnostic misclassification.

Alzheimer Disease↗

LRTae: improving statistical power for genetic association with case/control data when phenotype and/or genotype misclassification errors are present.

BACKGROUND: In the field of statistical genetics, phenotype and genotype misclassification errors can substantially reduce power to detect association with genetic case/control studies. Misclassification also can bias population frequency parameters such as genotype, haplotype, or multi-locus genotype frequencies. These problems are of particular concern in case/control designs because, short of repeated sampling, there is no way to detect misclassification errors. We developed a double-sampling procedure for case/control genetic association using a likelihood ratio test framework. Different approaches have been proposed to deal with misclassification errors. We have chosen the likelihood framework because of the ease with which misclassification probabilities may be incorporated into in the statistical framework and hypothesis testing. The statistic is called the Likelihood Ratio Test allowing for errors (LRTae) and is freely available via software download. RESULTS: We applied our procedure to 10,000 replicates of simulated case/control data in which we introduced phenotype misclassification errors. The phenotype considered is Ankylosing Spondylitis (AS). The LRTae method power was always greater than LRTstd power for the significance levels considered (5%, 1%, 0.1%, 0.01%). Power gains for the LRTae method over the LRTstd method increased as the significance level became more stringent. Multi-locus genotype frequency estimates using LRTae method were more accurate than estimates using LRTstd method. CONCLUSION: The LRTae method can be applied to single-locus genotypes, multi-locus genotypes, or multi-locus haplotypes in a case/control framework and can be more powerful to detect association in case/control studies when both genotype and/or phenotype errors are present. Furthermore, the LRTae method provides asymptotically unbiased estimates of case and control genotype frequencies, as well as estimates of phenotype and/or genotype misclassification rates.

Case-Control Studies↗

Serine/arginine-rich splicing factors belong to a class of intrinsically disordered proteins.

Serine/arginine-rich (SR) splicing factors play an important role in constitutive and alternative splicing as well as during several steps of RNA metabolism. Despite the wealth of functional information about SR proteins accumulated to-date, structural knowledge about the members of this family is very limited. To gain a better insight into structure-function relationships of SR proteins, we performed extensive sequence analysis of SR protein family members and combined it with ordered/disordered structure predictions. We found that SR proteins have properties characteristic of intrinsically disordered (ID) proteins. The amino acid composition and sequence complexity of SR proteins were very similar to those of the disordered protein regions. More detailed analysis showed that the SR proteins, and their RS domains in particular, are enriched in the disorder-promoting residues and are depleted in the order-promoting residues as compared to the entire human proteome. Moreover, disorder predictions indicated that RS domains of SR proteins were completely unstructured. Two different classification methods, the charge-hydropathy measure and the cumulative distribution function (CDF) of the disorder scores, were in agreement with each other, and they both strongly predicted members of the SR protein family to be disordered. This study emphasizes the importance of the disordered structure for several functions of SR proteins, such as for spliceosome assembly and for interaction with multiple partners. In addition, it demonstrates the usefulness of order/disorder predictions for inferring protein structure from sequence.

Amino Acids↗

Computing asymptotic power and sample size for case-control genetic association studies in the presence of phenotype and/or genotype misclassification errors.

It is well established that phenotype and genotype misclassification errors reduce the power to detect genetic association. Resampling a subset of the data (e.g, double-sampling) of genotype and/or phenotype with a gold standard measurement is one method to address this issue. We derive the non-centrality parameter (NCP) for the recently published Likelihood Ratio Test Allowing for Error (LRTae) in the presence of random phenotype and genotype errors. With the NCP, power and sample size can be analytically determined at any significance level. We verify analytic power with simulations using a 2**k factorial design given high and low settings of: case and control genotype frequencies, phenotype and genotype misclassification probabilities, total sample size, ratio of cases to controls, and proportions of phenotype and/or genotype double-samples. We also perform example applications of our method assuming equal costs for the LRTae method and the standard method that does not use double-sample information (LRTstd) to determine if power gain due to double-sampling a proportion of samples outweighs the reduction in sample size due to additional costs in obtaining double-samples. Our results showed a median difference of at most 0.01 between analytic and simulation power for the factorial design settings, with maximum difference of 0.054. For our cost/benefits analysis calculations, results for genotype errors are that double-sampling appears most beneficial (in terms of power gain) when cost of double-sampling is relatively low, irrespective of the proportion of individuals double-sampled. In the presence of phenotype error, there is always power gain using the LRTae method for the parameter settings considered. We have freely available software that performs power and sample size calculations for the LRTae method and cost/benefits analyses comparing power for LRTae and LRTstd methods assuming equal costs.

Journal Article↗

Expanded genome scan in extended families with age-related macular degeneration.

PURPOSE: To investigate further the genetic contribution to age-related macular degeneration (AMD), increasing the power of a previous analysis and reproducing the original findings. METHODS: A large cohort of families with this condition was assembled, and an expanded genome scan was performed with 556 microsatellite markers. In 2003, the results were reported of a genome-wide linkage analysis of 70 of these pedigrees. Members of 51 new families have now been ascertained and many of the original pedigrees expanded. Parametric and nonparametric linkage analyses were performed with a denser map of markers. In addition, analyses were performed with the sample stratified by age at ascertainment and by two major advanced phenotypes for the disease: neovascular AMD (choroidal neovascularization) and geographic atrophy. RESULTS: The results corroborate the macular degeneration-susceptibility loci consistently reported by the authors and others in genome-wide scans. New loci were identified, including the finding of a two-point HLOD of 3.70 at 6q25.2. CONCLUSIONS: The results suggest that the use of families enriched in predisposition to AMD has legitimacy. Genetic analyses of a genome-wide scan performed on our large cohort of families add further confirmatory evidence that susceptibility loci lie on 1q, 3p, 9q, and 10q. Furthermore, new loci have been identified, including a locus on 6q.

Aged↗

Precision and type I error rate in the presence of genotype errors and missing parental data: a comparison between the original transmission disequilibrium test (TDT) and TDTae statistics.

BACKGROUND: Two factors impacting robustness of the original transmission disequilibrium test (TDT) are: i) missing parental genotypes and ii) undetected genotype errors. While it is known that independently these factors can inflate false-positive rates for the original TDT, no study has considered either the joint impact of these factors on false-positive rates or the precision score of TDT statistics regarding these factors. By precision score, we mean the absolute difference between disease gene position and the position of markers whose TDT statistic exceeds some threshold. METHODS: We apply our transmission disequilibrium test allowing for errors (TDTae) and the original TDT to phenotype and modified single-nucleotide polymorphism genotype simulation data from Genetic Analysis Workshop. We modify genotype data by randomly introducing genotype errors and removing a percentage of parental genotype data. We compute empirical distributions of each statistic's precision score for a chromosome harboring a simulated disease locus. We also consider inflation in type I error by studying markers on a chromosome harboring no disease locus. RESULTS: The TDTae shows median precision scores of approximately 13 cM, 2 cM, 0 cM, and 0 cM at the 5%, 1%, 0.1%, and 0.01% significance levels, respectively. By contrast, the original TDT shows median precision scores of approximately 23 cM, 21 cM, 15 cM, and 7 cM at the corresponding significance levels, respectively. For null chromosomes, the original TDT falsely rejects the null hypothesis for 28.8%, 14.8%, 5.4%, and 1.7% at the 5%, 1%, 0.1% and 0.01%, significance levels, respectively, while TDTae maintains the correct false-positive rate. CONCLUSION: Because missing parental genotypes and undetected genotype errors are unknown to the investigator, but are expected to be increasingly prevalent in multilocus datasets, we strongly recommend TDTae methods as a standard procedure, particularly where stricter significance levels are required.

Chromosomes, Human, Pair 3↗

PAWE-3D: visualizing power for association with error in case-control genetic studies of complex traits.

UNLABELLED: A website that plots power and sample size calculations over a range of up to eight parameters (including diagnostic misclassification error parameters) for two commonly used statistical tests of genetic association, the linear trend test and the genotypic test of association. AVAILABILITY: This method is made available via the website http://linkage.rockefeller.edu/pawe3d/ CONTACT: pawe3d@linkage.rockefeller.edu.

Algorithms↗

Power and sample size calculations in the presence of phenotype errors for case/control genetic association studies.

BACKGROUND: Phenotype error causes reduction in power to detect genetic association. We present a quantification of phenotype error, also known as diagnostic error, on power and sample size calculations for case-control genetic association studies between a marker locus and a disease phenotype. We consider the classic Pearson chi-square test for independence as our test of genetic association. To determine asymptotic power analytically, we compute the distribution's non-centrality parameter, which is a function of the case and control sample sizes, genotype frequencies, disease prevalence, and phenotype misclassification probabilities. We derive the non-centrality parameter in the presence of phenotype errors and equivalent formulas for misclassification cost (the percentage increase in minimum sample size needed to maintain constant asymptotic power at a fixed significance level for each percentage increase in a given misclassification parameter). We use a linear Taylor Series approximation for the cost of phenotype misclassification to determine lower bounds for the relative costs of misclassifying a true affected (respectively, unaffected) as a control (respectively, case). Power is verified by computer simulation. RESULTS: Our major findings are that: (i) the median absolute difference between analytic power with our method and simulation power was 0.001 and the absolute difference was no larger than 0.011; (ii) as the disease prevalence approaches 0, the cost of misclassifying a unaffected as a case becomes infinitely large while the cost of misclassifying an affected as a control approaches 0. CONCLUSION: Our work enables researchers to specifically quantify power loss and minimum sample size requirements in the presence of phenotype errors, thereby allowing for more realistic study design. For most diseases of current interest, verifying that cases are correctly classified is of paramount importance.

Alzheimer Disease↗

Complement factor H polymorphism in age-related macular degeneration.

Age-related macular degeneration (AMD) is a major cause of blindness in the elderly. We report a genome-wide screen of 96 cases and 50 controls for polymorphisms associated with AMD. Among 116,204 single-nucleotide polymorphisms genotyped, an intronic and common variant in the complement factor H gene (CFH) is strongly associated with AMD (nominal P value <10(-7)). In individuals homozygous for the risk allele, the likelihood of AMD is increased by a factor of 7.4 (95% confidence interval 2.9 to 19). Resequencing revealed a polymorphism in linkage disequilibrium with the risk allele representing a tyrosine-histidine change at amino acid 402. This polymorphism is in a region of CFH that binds heparin and C-reactive protein. The CFH gene is located on chromosome 1 in a region repeatedly linked to AMD in family-based studies.

Aged↗

Increasing power for tests of genetic association in the presence of phenotype and/or genotype error by use of double-sampling.

Phenotype and/or genotype misclassification can: significantly increase type II error probabilities for genetic case/control association, causing decrease in statistical power; and produce inaccurate estimates of population frequency parameters. We present a method, the likelihood ratio test allowing for errors (LRTae) that incorporates double-sample information for phenotypes and/or genotypes on a sub-sample of cases/controls. Population frequency parameters and misclassification probabilities are determined using a double-sample procedure as implemented in the Expectation-Maximization (EM) method. We perform null simulations assuming a SNP marker or a 4-allele (multi-allele) marker locus. To compare our method with the standard method that makes no adjustment for errors (LRTstd), we perform power simulations using a 2/k factorial design with high and low settings of: case/control samples, phenotype/genotype costs, double-sampled phenotypes/genotypes costs, phenotype/genotype error, and proportions of double-sampled individuals. All power simulations are performed fixing equal costs for the LRTstd and LRTae methods. We also consider case/control ApoE genotype data for an actual Alzheimer's study. The LRTae method maintains correct type I error proportions for all null simulations and all significance level thresholds (10%, 5%, 1%). LRTae average estimates of population frequencies and misclassification probabilities are equal to the true values, with variances of 10e-7 to 10e-8. For power simulations, the median power difference LRTae-LRTstd at the 5% significance level is 0.06 for multi-allele data and 0.01 for SNP data. For the ApoE data example, the LRTae and LRTstd p-values are 5.8 x 10e-5 and 1.6 x 10e-3, respectively. The increase in significance is due to adjustment in the LRTae for misclassification of the most commonly reported risk allele. We have developed freely available software that performs our LRTae statistic.

Journal Article↗

A transmission disequilibrium test for general pedigrees that is robust to the presence of random genotyping errors and any number of untyped parents.

Two issues regarding the robustness of the original transmission disequilibrium test (TDT) developed by Spielman et al are: (i) missing parental genotype data and (ii) the presence of undetected genotype errors. While extensions of the TDT that are robust to items (i) and (ii) have been developed, there is to date no single TDT statistic that is robust to both for general pedigrees. We present here a likelihood method, the TDT(ae), which is robust to these issues in general pedigrees. The TDT(ae) assumes a more general disease model than the traditional TDT, which assumes a multiplicative inheritance model for genotypic relative risk. Our model is based on Weinberg's work. To assess robustness, we perform simulations. Also, we apply our method to two data sets from actual diseases: psoriasis and sitosterolemia. Maximization under alternative and null hypotheses is performed using Powell's method. Results of our simulations indicate that our method maintains correct type I error rates at the 1, 5, and 10% levels of significance. Furthermore, a Kolmorogov-Smirnoff Goodness of Fit test suggests that the data are drawn from a central chi2 with 2 df, the correct asymptotic null distribution. The psoriasis results suggest two loci as being significantly linked to the disease, even in the presence of genotyping errors and missing data, and the sitosterolemia results show a P-value of 1.5 x 10(-9) for the marker locus nearest to the sitosterolemia disease genes. We have developed software to perform TDT(ae) calculations, which may be accessed from our ftp site.

Computer Simulation↗

Quantifying the percent increase in minimum sample size for SNP genotyping errors in genetic model-based association studies.

Kang et al. [Genet Epidemiol 2004;26:132-141] addressed the question of which genotype misclassification errors are most costly, in terms of minimum percentage increase in sample size necessary (%MSSN) to maintain constant asymptotic power and significance level, when performing case/control studies of genetic association in a genetic model-free setting. They answered the question for single nucleotide polymorphisms (SNPs) using the 2 x 3 chi2 test of independence. We address the same question here for a genetic model-based framework. The genetic model parameters considered are: disease model (dominant, recessive), genotypic relative risk, SNP (marker) and disease allele frequency, and linkage disequilibrium. %MSSN coefficients of each of the six possible error rates are determined by expanding the non-centrality parameter of the asymptotic distribution of the 2 x 3 chi2 test under a specified alternative hypothesis to approximate %MSSN using a linear Taylor series in the error rates. In this work we assume errors misclassifying one homozygote as another homozygote are 0, since these errors are thought to rarely occur in practice. Our findings are that there are settings of the genetic model parameters that lead to large total %MSSN for both dominant and recessive models. As SNP minor allele approaches 0, total %MSSN increases without bound, independent of other genetic model parameters. In general, %MSSN is a complex function of the genetic model parameters. Use of SNPs with small minor allele frequency requires careful attention to frequency of genotyping errors to insure that power specifications are met. Software to perform these calculations for study design is available, and an example of its use to study a disease is given.

Alleles↗