PubMed Health⌕ Search

Biomedical subjects

Derek Gordon

Publications and source records attributed to Derek Gordon.

At least 19 recordsLinked to original sources

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1. Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows that Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occur alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD.

Animals↗

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

UNLABELLED: Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1 . Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occurs alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD. SIGNIFICANCE STATEMENT: We generated mouse models that express mutations in high-confidence genes linked to Tourette disorder (TD). These models show sensorimotor and cognitive behavioral phenotypes resembling TD-like behaviors. Sensorimotor gating deficits and repetitive motor behaviors are attenuated by drugs that act on dopamine. Reward learning and striatal dopamine is enhanced. Brain development is grossly normal, including cortical layering and patterning of major axon tracts. Further, no signs of striatal interneuron loss are detected. Interestingly, behavioral phenotypes in affected females can be more pronounced than in males, despite male sex bias in the diagnosis of TD. These novel mouse models with construct, face, and predictive validity provide a new resource to study neural substrates that cause tics and related behavioral phenotypes in TD.

Preprint↗

The effects of SNP genotyping errors on the power of the Cochran-Armitage linear trend test for case/control association studies.

The questions addressed in this paper are: What single nucleotide polymorphism (SNP) genotyping errors are most costly, in terms of minimum sample size necessary (MSSN) to maintain constant asymptotic power and significance level, when performing case-control studies of genetic association applying the Cochran-Armitage trend test? And which trend test or chi2 test is more powerful under standard genetic models with genotyping errors? Our strategy is to expand the non-centrality parameter of the asymptotic distribution of the trend test to approximate the MSSN using a Taylor series linear in the genotyping error rates. We apply our strategy to example scenarios that assume recessive, dominant, additive, or over-dominant disease models. The most costly errors are recording the more common homozygote as the less common homozygote, and the more common homozygote as the heterozygote, with MSSN that become indefinitely large as the minor SNP allele frequency approaches zero. Misclassifying the heterozygote as the less common homozygote is costly when using the recessive trend test on data from a recessive model. The chi2 test has power close to, but less than, the optimal trend test and is never dominated over all genetic models studied by any specific trend test.

Case-Control Studies↗

Are molecular haplotypes worth the time and expense? A cost-effective method for applying molecular haplotypes.

Because current molecular haplotyping methods are expensive and not amenable to automation, many researchers rely on statistical methods to infer haplotype pairs from multilocus genotypes, and subsequently treat these inferred haplotype pairs as observations. These procedures are prone to haplotype misclassification. We examine the effect of these misclassification errors on the false-positive rate and power for two association tests. These tests include the standard likelihood ratio test (LRTstd) and a likelihood ratio test that employs a double-sampling approach to allow for the misclassification inherent in the haplotype inference procedure (LRTae). We aim to determine the cost-benefit relationship of increasing the proportion of individuals with molecular haplotype measurements in addition to genotypes to raise the power gain of the LRTae over the LRTstd. This analysis should provide a guideline for determining the minimum number of molecular haplotypes required for desired power. Our simulations under the null hypothesis of equal haplotype frequencies in cases and controls indicate that (1) for each statistic, permutation methods maintain the correct type I error; (2) specific multilocus genotypes that are misclassified as the incorrect haplotype pair are consistently misclassified throughout each entire dataset; and (3) our simulations under the alternative hypothesis showed a significant power gain for the LRTae over the LRTstd for a subset of the parameter settings. Permutation methods should be used exclusively to determine significance for each statistic. For fixed cost, the power gain of the LRTae over the LRTstd varied depending on the relative costs of genotyping, molecular haplotyping, and phenotyping. The LRTae showed the greatest benefit over the LRTstd when the cost of phenotyping was very high relative to the cost of genotyping. This situation is likely to occur in a replication study as opposed to a whole-genome association study.

Cost-Benefit Analysis↗

Increase in linkage information by stratification of pedigree data into gold-standard and standard diagnoses: application to the NIMH Alzheimer Disease Genetics Initiative Dataset.

Patients diagnosed with a standard clinical method (subject to misclassification error) are often combined with patients diagnosed with a gold-standard method (with zero or very small misclassification error) in family-based studies of complex disease. For example, non-autopsied patients (NAP) are often included along with autopsy-proven (AP) patients in family-based studies of complex diseases, such as Alzheimer's disease (AD). Theoretical and simulation studies suggest that certain misclassification errors can result in severe reduction of power in genetic linkage and association analyses and that phenotype (or diagnostic) error can produce misleading results. Morton's test for heterogeneity can identify genomic regions where error may have led to loss in power. We applied this test to pedigree data from the NIMH Alzheimer's Disease Genetics Initiative Database separated into AP and NAP pedigrees. Morton's test identified one highly significant region of heterogeneity on chromosome 2. The source of the heterogeneity was due to significant indication of linkage in the AP pedigrees at position 109 cM (p value = 6.68 x 10(-5)) with no indication in the NAP pedigrees. Furthermore, Morton's test showed no evidence for heterogeneity on chromosome 19 in early-onset pedigrees that showed highly significant evidence for linkage in other published reports. These results suggest that supplementing linkage analysis with Morton's test can be usefully applied to genetic data sets that have AP and NAP samples, or other sample mixtures that include a 'gold standard' subgroup with reduced error rate, to increase power to detect linkage in the presence of diagnostic misclassification.

Alzheimer Disease↗

LRTae: improving statistical power for genetic association with case/control data when phenotype and/or genotype misclassification errors are present.

BACKGROUND: In the field of statistical genetics, phenotype and genotype misclassification errors can substantially reduce power to detect association with genetic case/control studies. Misclassification also can bias population frequency parameters such as genotype, haplotype, or multi-locus genotype frequencies. These problems are of particular concern in case/control designs because, short of repeated sampling, there is no way to detect misclassification errors. We developed a double-sampling procedure for case/control genetic association using a likelihood ratio test framework. Different approaches have been proposed to deal with misclassification errors. We have chosen the likelihood framework because of the ease with which misclassification probabilities may be incorporated into in the statistical framework and hypothesis testing. The statistic is called the Likelihood Ratio Test allowing for errors (LRTae) and is freely available via software download. RESULTS: We applied our procedure to 10,000 replicates of simulated case/control data in which we introduced phenotype misclassification errors. The phenotype considered is Ankylosing Spondylitis (AS). The LRTae method power was always greater than LRTstd power for the significance levels considered (5%, 1%, 0.1%, 0.01%). Power gains for the LRTae method over the LRTstd method increased as the significance level became more stringent. Multi-locus genotype frequency estimates using LRTae method were more accurate than estimates using LRTstd method. CONCLUSION: The LRTae method can be applied to single-locus genotypes, multi-locus genotypes, or multi-locus haplotypes in a case/control framework and can be more powerful to detect association in case/control studies when both genotype and/or phenotype errors are present. Furthermore, the LRTae method provides asymptotically unbiased estimates of case and control genotype frequencies, as well as estimates of phenotype and/or genotype misclassification rates.

Case-Control Studies↗

A detailed Hapmap of the Sitosterolemia locus spanning 69 kb; differences between Caucasians and African-Americans.

BACKGROUND: Sitosterolemia is an autosomal recessive disorder that maps to the sitosterolemia locus, STSL, on human chromosome 2p21. Two genes, ABCG5 and ABCG8, comprise the STSL and mutations in either cause sitosterolemia. ABCG5 and ABCG8 are thought to have evolved by gene duplication event and are arranged in a head-to-head configuration. We report here a detailed characterization of the STSL in Caucasian and African-American cohorts. METHODS: Caucasian and African-American DNA samples were genotypes for polymorphisms at the STSL locus and haplotype structures determined for this locus RESULTS: In the Caucasian population, 13 variant single nucleotide polymorphisms (SNPs) were identified and resulting in 24 different haplotypes, compared to 11 SNPs in African-Americans resulting in 40 haplotypes. Three polymorphisms in ABCG8 were unique to the Caucasian population (E238L, INT10-50 and G575R), whereas one variant (A259V) was unique to the African-American population. Allele frequencies of SNPs varied also between these populations. CONCLUSION: We confirmed that despite their close proximity to each other, significantly more variations are present in ABCG8 compared to ABCG5. Pairwise D' values showed wide ranges of variation, indicating some of the SNPs were in strong linkage disequilibrium (LD) and some were not. LD was more prevalent in Caucasians than in African-Americans, as would be expected. These data will be useful in analyzing the proposed role of STSL in processes ranging from responsiveness to cholesterol-lowering drugs to selective sterol absorption.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

Computing asymptotic power and sample size for case-control genetic association studies in the presence of phenotype and/or genotype misclassification errors.

It is well established that phenotype and genotype misclassification errors reduce the power to detect genetic association. Resampling a subset of the data (e.g, double-sampling) of genotype and/or phenotype with a gold standard measurement is one method to address this issue. We derive the non-centrality parameter (NCP) for the recently published Likelihood Ratio Test Allowing for Error (LRTae) in the presence of random phenotype and genotype errors. With the NCP, power and sample size can be analytically determined at any significance level. We verify analytic power with simulations using a 2**k factorial design given high and low settings of: case and control genotype frequencies, phenotype and genotype misclassification probabilities, total sample size, ratio of cases to controls, and proportions of phenotype and/or genotype double-samples. We also perform example applications of our method assuming equal costs for the LRTae method and the standard method that does not use double-sample information (LRTstd) to determine if power gain due to double-sampling a proportion of samples outweighs the reduction in sample size due to additional costs in obtaining double-samples. Our results showed a median difference of at most 0.01 between analytic and simulation power for the factorial design settings, with maximum difference of 0.054. For our cost/benefits analysis calculations, results for genotype errors are that double-sampling appears most beneficial (in terms of power gain) when cost of double-sampling is relatively low, irrespective of the proportion of individuals double-sampled. In the presence of phenotype error, there is always power gain using the LRTae method for the parameter settings considered. We have freely available software that performs power and sample size calculations for the LRTae method and cost/benefits analyses comparing power for LRTae and LRTstd methods assuming equal costs.

Journal Article↗

Localization of breast cancer susceptibility loci by genome-wide SNP linkage disequilibrium mapping.

We studied the feasibility of a novel approach to localize breast cancer susceptibility genes, using a low-density genome-wide panel of single-nucleotide polymorphisms and taking advantage of large regions of linkage disequilibrium (LD) flanking Jewish disease genes in high-risk cases. With Affymetrix GeneChip arrays, we genotyped 8,576 polymorphisms in three sets of Ashkenazi Jewish breast cancer cases: a "validation" set of 27 breast cancer cases, all of whom carried the BRCA2*6174delT founder mutation; a "field" set of 19 breast cancer cases from male breast cancer kindreds, which simulated conditions for finding new genes; and a "test" set of 57 probands from breast cancer kindreds (4 or more cases/kindred), in which mutations in BRCA1 and BRCA2 had been excluded. To identify associations, we compared the frequency of genotypes and haplotypes in cases vs. controls by the Fisher's exact test and a maximum likelihood ratio test. In the "validation" set, we demonstrated the presence of a region of linkage disequilibrium on BRCA2*6174delT chromosomes that spanned over 5 million bases. In the "field" set, we showed that this large region of linkage disequilibrium flanking BRCA2 was detectable despite the presence of heterogeneity in the sample set. Finally, in the "test" set, at least three regions of interest emerged that could contain novel breast cancer genes, one of which had been identified previously by linkage analysis. While these results demonstrate the feasibility of genome-wide association strategies, further application of this approach will critically depend on optimizing the density and distribution of SNPs and the size and type of study design.

Adult↗

Association analysis of polymorphisms in serotonin 1B receptor (HTR1B) gene with heroin addiction: a comparison of molecular and statistically estimated haplotypes.

OBJECTIVES: 5-Hydroxytryptamine (serotonin)-1B receptors (HTR1B) may play an important role in psychiatric disorders and drug and alcohol dependence. In this study we report on genotype, molecular haplotype and statistically estimated haplotype analyses of previously identified polymorphisms in positions -261T>G, -161A>T, 129C>T, 861G>C and 1180A>G of the HTR1B gene in ethnically diverse populations (African-Americans, Caucasians, Hispanics and Asians) including 235 former heroin addicts and 161 control subjects from New York City. The objectives were to test for an association of molecular and statistically estimated haplotypes and genotypes in HTR1B gene with heroin addiction and to compare results provided by molecular and statistically estimated haplotyping methods. METHODS: Genotype analysis was performed using a standard TaqMan protocol. Molecular haplotype analysis of the subset of polymorphisms consisting of -261T>G, -161A>T and 129C>T was performed using a protocol specially designed by our group, using fluorescent PCR. This is based on use of allele-specific primers complementary to flanking polymorphisms and a fluorescently labeled sequence-specific TaqMan probe set complementary to an internal polymorphism of the haplotype region. Every individual's statistically inferred haplotype pair agreed with the individual's haplotype pair determined by molecular haplotyping. RESULTS AND CONCLUSION: A point-wise significant association of haplotype pairs containing allele G at position 1180 with protective effect from heroin addiction in Caucasians was found. A point-wise nominally significant association of allele 1180G with a protective effect from heroin addiction was found in Caucasians. Statistically significant differences across four ethnic groups in control subjects for allelic frequencies of -261T>G and -161A>T were found.

Black or African American↗

Precision and type I error rate in the presence of genotype errors and missing parental data: a comparison between the original transmission disequilibrium test (TDT) and TDTae statistics.

BACKGROUND: Two factors impacting robustness of the original transmission disequilibrium test (TDT) are: i) missing parental genotypes and ii) undetected genotype errors. While it is known that independently these factors can inflate false-positive rates for the original TDT, no study has considered either the joint impact of these factors on false-positive rates or the precision score of TDT statistics regarding these factors. By precision score, we mean the absolute difference between disease gene position and the position of markers whose TDT statistic exceeds some threshold. METHODS: We apply our transmission disequilibrium test allowing for errors (TDTae) and the original TDT to phenotype and modified single-nucleotide polymorphism genotype simulation data from Genetic Analysis Workshop. We modify genotype data by randomly introducing genotype errors and removing a percentage of parental genotype data. We compute empirical distributions of each statistic's precision score for a chromosome harboring a simulated disease locus. We also consider inflation in type I error by studying markers on a chromosome harboring no disease locus. RESULTS: The TDTae shows median precision scores of approximately 13 cM, 2 cM, 0 cM, and 0 cM at the 5%, 1%, 0.1%, and 0.01% significance levels, respectively. By contrast, the original TDT shows median precision scores of approximately 23 cM, 21 cM, 15 cM, and 7 cM at the corresponding significance levels, respectively. For null chromosomes, the original TDT falsely rejects the null hypothesis for 28.8%, 14.8%, 5.4%, and 1.7% at the 5%, 1%, 0.1% and 0.01%, significance levels, respectively, while TDTae maintains the correct false-positive rate. CONCLUSION: Because missing parental genotypes and undetected genotype errors are unknown to the investigator, but are expected to be increasingly prevalent in multilocus datasets, we strongly recommend TDTae methods as a standard procedure, particularly where stricter significance levels are required.

Chromosomes, Human, Pair 3↗

Characteristics of replicated single-nucleotide polymorphism genotypes from COGA: Affymetrix and Center for Inherited Disease Research.

Genetic Analysis Workshop 14 provided re-genotyped single-nucleotide polymorphism (SNP) data. Specifically, both Center for Inherited Disease Research (CIDR) and Affymetrix genotyped the same 11,560 SNPs from the Affymetrix GeneChip Mapping 10K Array marker set on the same 184 individuals from the Collaborative Study on the Genetics of Alcoholism database. While the inconsistency rate between CIDR and Affymetrix (two different genotypes for the same subject) was low (0.2%), the non-replication rate (two different genotypes for the same subject or one identified genotype and one missing genotype) was substantial (9.5%). The missing data could be from no-call regions, which is inconsistent with recent recommendations about the use of no-call regions in association tests. In addition, no-call regions would suggest that the actual inconsistency rate is higher than reported. A high inconsistency rate has significant impact on power in related hypothesis tests. In addition, the data are consistent with assumptions made in a recently proposed likelihood ratio test of association for re-genotyped data.

Alcoholism↗

Disease relevant HLA class II alleles isolated by genotypic, haplotypic, and sequence analysis in North American Caucasians with pemphigus vulgaris.

Early studies of genetic susceptibility to pemphigus vulgaris (PV) showed associations between human leukocyte antigen (HLA) DR4 and DR6 and disease. The emergence of DNA sequencing techniques has implicated numerous DRB1 and DQB1 loci in various populations, leading to confusion regarding which exact alleles confer susceptibility. The strong linkage disequilibrium among DR and DQ HLA alleles further complicates the investigation of the true susceptibility loci. In this study, we report genotyping data for the largest sampling of North American Caucasian non-Jewish and Ashkenazi Jewish PV patients studied to date and compare our data with other population studies. To pinpoint true susceptibility, alleles among overrepresented sequences, we applied a step-wise reductionist analysis through (1) determination of the degree of linkage disequilibrium (LD) between purportedly associated alleles, (2) haplotype frequencies comparisons, and (3) primary sequence comparisons of disease-associated versus non-disease-associated alleles to identify crucial differences in amino acid residues in putative peptide binding pockets. Collectively, our data provide extended support for the hypothesis that the HLA associations in Caucasian PV patients map to DRB1*0402 and DQB1*0503 alone. Further structure-function studies will be required to define the exact mechanisms of HLA-mediated control of susceptibility and resistance to disease.

Adult↗

Localization of PSORS1 to a haplotype block harboring HLA-C and distinct from corneodesmosin and HCR.

Psoriasis is a complex inflammatory disease of the skin affecting 1-2% of the Caucasian population. Associations with alleles from the HLA class I region (now known as PSORS1), particularly HLA-Cw*0602, were described over 20 years ago. However, extensive linkage disequilibrium (LD) within this region has made it difficult to identify the true susceptibility allele from this region. A variety of genes and regions from a 238-kb interval extending from HLA-B to corneodesmosin (CDSN) have been proposed to harbor PSORS1. In order to identify the minimum block of LD in the MHC class I region associated with psoriasis we performed a comprehensive case/control and family-based association study on 242 Northern European psoriasis families and two separate European control populations. High resolution HLA typing of HLA-A, -B and -C alleles was performed, in addition to the genotyping of 18 polymorphic microsatellites and 36 SNPs from a 772-kb segment of the HLA class I region harboring the previously described interval. This corresponded on average to one SNP every 7 kb in the candidate 238 kb region. With all tests, the association was the strongest with single markers and haplotypes from a block of LD harboring HLA-C and SNP n.9. Logistic regression analyses indicated that association seen with candidate genes from the interval such as CDSN and HCR was entirely dependent on association with HLA-Cw*0602 and SNP n.9-G alleles. The previously reported association with CDSN and HCR was observed to be due to the existence of the associated alleles lying on the most commonly over-transmitted haplotype. Rare over-transmitted haplotypes also harbored HLA-Cw*12 alleles. HLA-Cw*12 family members are closely related to HLA Cw*0602, sharing identical sequences in their alpha-2 domains, peptide-binding pockets A, D and E and all 3' introns. The introduction of a potential binding site for the RUNX/AML family of transcription factors in intron 7, is also specific to these HLA-C alleles. These variants need to be investigated further for their role as PSORS1.

Case-Control Studies↗

Power and sample size calculations for genetic case/control studies using gene-centric SNP maps: application to human chromosomes 6, 21, and 22 in three populations.

Power and sample size calculations are critical parts of any research design for genetic association. We present a method that utilizes haplotype frequency information and average marker-marker linkage disequilibrium on SNPs typed in and around all genes on a chromosome. The test statistic used is the classic likelihood ratio test applied to haplotypes in case/control populations. Haplotype frequencies are computed through specification of genetic model parameters. Power is determined by computation of the test's non-centrality parameter. Power per gene is computed as a weighted average of the power assuming each haplotype is associated with the trait. We apply our method to genotype data from dense SNP maps across three entire chromosomes (6, 21, and 22) for three different human populations (African-American, Caucasian, Chinese), three different models of disease (additive, dominant, and multiplicative) and two trait allele frequencies (rare, common). We perform a regression analysis using these factors, average marker-marker disequilibrium, and the haplotype diversity across the gene region to determine which factors most significantly affect average power for a gene in our data. Also, as a 'proof of principle' calculation, we perform power and sample size calculations for all genes within 100 kb of the PSORS1 locus (chromosome 6) for a previously published association study of psoriasis. Results of our regression analysis indicate that four highly significant factors that determine average power to detect association are: disease model, average marker-marker disequilibrium, haplotype diversity, and the trait allele frequency. These findings may have important implications for the design of well-powered candidate gene association studies. Our power and sample size calculations for the PSORS1 gene appear consistent with published findings, namely that there is substantial power (>0.99) for most genes within 100 kb of the PSORS1 locus at the 0.01 significance level.

Black or African American↗

PAWE-3D: visualizing power for association with error in case-control genetic studies of complex traits.

UNLABELLED: A website that plots power and sample size calculations over a range of up to eight parameters (including diagnostic misclassification error parameters) for two commonly used statistical tests of genetic association, the linear trend test and the genotypic test of association. AVAILABILITY: This method is made available via the website http://linkage.rockefeller.edu/pawe3d/ CONTACT: pawe3d@linkage.rockefeller.edu.

Algorithms↗

Power and sample size calculations in the presence of phenotype errors for case/control genetic association studies.

BACKGROUND: Phenotype error causes reduction in power to detect genetic association. We present a quantification of phenotype error, also known as diagnostic error, on power and sample size calculations for case-control genetic association studies between a marker locus and a disease phenotype. We consider the classic Pearson chi-square test for independence as our test of genetic association. To determine asymptotic power analytically, we compute the distribution's non-centrality parameter, which is a function of the case and control sample sizes, genotype frequencies, disease prevalence, and phenotype misclassification probabilities. We derive the non-centrality parameter in the presence of phenotype errors and equivalent formulas for misclassification cost (the percentage increase in minimum sample size needed to maintain constant asymptotic power at a fixed significance level for each percentage increase in a given misclassification parameter). We use a linear Taylor Series approximation for the cost of phenotype misclassification to determine lower bounds for the relative costs of misclassifying a true affected (respectively, unaffected) as a control (respectively, case). Power is verified by computer simulation. RESULTS: Our major findings are that: (i) the median absolute difference between analytic power with our method and simulation power was 0.001 and the absolute difference was no larger than 0.011; (ii) as the disease prevalence approaches 0, the cost of misclassifying a unaffected as a case becomes infinitely large while the cost of misclassifying an affected as a control approaches 0. CONCLUSION: Our work enables researchers to specifically quantify power loss and minimum sample size requirements in the presence of phenotype errors, thereby allowing for more realistic study design. For most diseases of current interest, verifying that cases are correctly classified is of paramount importance.

Alzheimer Disease↗