PubMed Health⌕ Search

Biomedical subjects

Stephen J Finch

Publications and source records attributed to Stephen J Finch.

At least 19 recordsLinked to original sources

The effects of SNP genotyping errors on the power of the Cochran-Armitage linear trend test for case/control association studies.

The questions addressed in this paper are: What single nucleotide polymorphism (SNP) genotyping errors are most costly, in terms of minimum sample size necessary (MSSN) to maintain constant asymptotic power and significance level, when performing case-control studies of genetic association applying the Cochran-Armitage trend test? And which trend test or chi2 test is more powerful under standard genetic models with genotyping errors? Our strategy is to expand the non-centrality parameter of the asymptotic distribution of the trend test to approximate the MSSN using a Taylor series linear in the genotyping error rates. We apply our strategy to example scenarios that assume recessive, dominant, additive, or over-dominant disease models. The most costly errors are recording the more common homozygote as the less common homozygote, and the more common homozygote as the heterozygote, with MSSN that become indefinitely large as the minor SNP allele frequency approaches zero. Misclassifying the heterozygote as the less common homozygote is costly when using the recessive trend test on data from a recessive model. The chi2 test has power close to, but less than, the optimal trend test and is never dominated over all genetic models studied by any specific trend test.

Case-Control Studies↗

Increase in linkage information by stratification of pedigree data into gold-standard and standard diagnoses: application to the NIMH Alzheimer Disease Genetics Initiative Dataset.

Patients diagnosed with a standard clinical method (subject to misclassification error) are often combined with patients diagnosed with a gold-standard method (with zero or very small misclassification error) in family-based studies of complex disease. For example, non-autopsied patients (NAP) are often included along with autopsy-proven (AP) patients in family-based studies of complex diseases, such as Alzheimer's disease (AD). Theoretical and simulation studies suggest that certain misclassification errors can result in severe reduction of power in genetic linkage and association analyses and that phenotype (or diagnostic) error can produce misleading results. Morton's test for heterogeneity can identify genomic regions where error may have led to loss in power. We applied this test to pedigree data from the NIMH Alzheimer's Disease Genetics Initiative Database separated into AP and NAP pedigrees. Morton's test identified one highly significant region of heterogeneity on chromosome 2. The source of the heterogeneity was due to significant indication of linkage in the AP pedigrees at position 109 cM (p value = 6.68 x 10(-5)) with no indication in the NAP pedigrees. Furthermore, Morton's test showed no evidence for heterogeneity on chromosome 19 in early-onset pedigrees that showed highly significant evidence for linkage in other published reports. These results suggest that supplementing linkage analysis with Morton's test can be usefully applied to genetic data sets that have AP and NAP samples, or other sample mixtures that include a 'gold standard' subgroup with reduced error rate, to increase power to detect linkage in the presence of diagnostic misclassification.

Alzheimer Disease↗

Computing asymptotic power and sample size for case-control genetic association studies in the presence of phenotype and/or genotype misclassification errors.

It is well established that phenotype and genotype misclassification errors reduce the power to detect genetic association. Resampling a subset of the data (e.g, double-sampling) of genotype and/or phenotype with a gold standard measurement is one method to address this issue. We derive the non-centrality parameter (NCP) for the recently published Likelihood Ratio Test Allowing for Error (LRTae) in the presence of random phenotype and genotype errors. With the NCP, power and sample size can be analytically determined at any significance level. We verify analytic power with simulations using a 2**k factorial design given high and low settings of: case and control genotype frequencies, phenotype and genotype misclassification probabilities, total sample size, ratio of cases to controls, and proportions of phenotype and/or genotype double-samples. We also perform example applications of our method assuming equal costs for the LRTae method and the standard method that does not use double-sample information (LRTstd) to determine if power gain due to double-sampling a proportion of samples outweighs the reduction in sample size due to additional costs in obtaining double-samples. Our results showed a median difference of at most 0.01 between analytic and simulation power for the factorial design settings, with maximum difference of 0.054. For our cost/benefits analysis calculations, results for genotype errors are that double-sampling appears most beneficial (in terms of power gain) when cost of double-sampling is relatively low, irrespective of the proportion of individuals double-sampled. In the presence of phenotype error, there is always power gain using the LRTae method for the parameter settings considered. We have freely available software that performs power and sample size calculations for the LRTae method and cost/benefits analyses comparing power for LRTae and LRTstd methods assuming equal costs.

Journal Article↗

The relationship of personality and behavioral development from adolescence to young adulthood and subsequent parenting behavior.

The purpose of the study was to examine the association of parental personality, behavior, and substance use during adolescence and adulthood as related to the later parent-offspring relationship. The sample consisted of 297 parents (M age 32 yr.), who were first interviewed at earlier points in their lives in childhood and early adolescence at six points in time, extending from 1983 to 2002. Multiple regression models showed that parents with certain earlier personality and behavioral attributes, e.g., more rebelliousness and more frequent tobacco use, had a more difficult relationship with their children. Findings indicated an association between the cumulative number of psychosocial risk factors in the parents and difficulties in the parent-child relationship. The findings suggested that interventions designed to decrease youths' substance abuse may increase the likelihood that later when they are parents they will form nurturing relationships with their children.

Adolescent↗

Smoking involvement during adolescence among African Americans and Puerto Ricans: risks to psychological and physical well-being in young adulthood.

The major aim of this study was to examine the longitudinal association between adolescent smoking involvement and self-reported psychological and physical outcomes in young adulthood. Participants included 333 African Americans and 329 Puerto Ricans who were surveyed in 1990 in their New York City schools and interviewed in 1995 and 2000-2001, primarily in their homes. The psychological outcomes included ego integration, symptoms of depression, anxiety, and interpersonal difficulty. The physical health measures included a general health rating, number of illnesses, and symptoms of ill health. Also, three scales measured problems due to alcohol, marijuana, and other illicit drug use. Smoking involvement varied by age, sex, and ethnicity but not by socioeconomic status nor by late adolescent parental status. Analysis showed that the relationships between adolescent smoking involvement and psychological and physical health problems in young adulthood remained significant even with control on demographic factors, earlier levels of the outcome variables, and marijuana use. The relationships between smoking behavior and problems with alcohol, marijuana and other illicit drug use were particularly strong. Thus, adolescent smoking seems to have a wide range of clinical implications for young adulthood.

Adaptation, Psychological↗

Characteristics of replicated single-nucleotide polymorphism genotypes from COGA: Affymetrix and Center for Inherited Disease Research.

Genetic Analysis Workshop 14 provided re-genotyped single-nucleotide polymorphism (SNP) data. Specifically, both Center for Inherited Disease Research (CIDR) and Affymetrix genotyped the same 11,560 SNPs from the Affymetrix GeneChip Mapping 10K Array marker set on the same 184 individuals from the Collaborative Study on the Genetics of Alcoholism database. While the inconsistency rate between CIDR and Affymetrix (two different genotypes for the same subject) was low (0.2%), the non-replication rate (two different genotypes for the same subject or one identified genotype and one missing genotype) was substantial (9.5%). The missing data could be from no-call regions, which is inconsistent with recent recommendations about the use of no-call regions in association tests. In addition, no-call regions would suggest that the actual inconsistency rate is higher than reported. A high inconsistency rate has significant impact on power in related hypothesis tests. In addition, the data are consistent with assumptions made in a recently proposed likelihood ratio test of association for re-genotyped data.

Alcoholism↗

A gene-model-free method for linkage analysis of a disease-related-trait based on analysis of proband/sibling pairs.

In this paper we investigate the power of finding linkage to a disease locus through analysis of the disease-related traits. We propose two family-based gene-model-free linkage statistics. Both involve considering the distribution of the number of alleles identical by descent with the proband and comparing siblings with the disease-related trait to those without the disease-related-trait. The objective is to find linkages to disease-related traits that are pleiotropic for both the disease and the disease-related-traits. The power of these statistics is investigated for Kofendrerd Personality Disorder-related traits a (Joining/founding cults) and trait b (Fear/discomfort with strangers) of the simulated data. The answers were known prior to the execution of the reported analyses. We find that both tests have very high power when applied to the samples created by combining the data of the three cities for which we have nuclear family data.

Chromosome Mapping↗

Using mixture models to characterize disease-related traits.

We consider 12 event-related potentials and one electroencephalogram measure as disease-related traits to compare alcohol-dependent individuals (cases) to unaffected individuals (controls). We use two approaches: 1) two-way analysis of variance (with sex and alcohol dependency as the factors), and 2) likelihood ratio tests comparing sex adjusted values of cases to controls assuming that within each group the trait has a 2 (or 3) component normal mixture distribution. In the second approach, we test the null hypothesis that the parameters of the mixtures are equal for the cases and controls. Based on the two-way analysis of variance, we find 1) males have significantly (p < 0.05) lower mean response values than females for 7 of these traits. 2) Alcohol-dependent cases have significantly lower mean response than controls for 3 traits. The mixture analysis of sex-adjusted values of 1 of these traits, the event-related potential obtained at the parietal midline channel (ttth4), found the appearance of a 3-component normal mixture in cases and controls. The mixtures differed in that the cases had significantly lower mean values than controls and significantly different mixing proportions in 2 of the 3 components. Implications of this study are: 1) Sex needs to be taken into account when studying risk factors for alcohol dependency to prevent finding a spurious association between alcohol dependency and the risk factor. 2) Mixture analysis indicates that for the event-related potential "ttth4", the difference observed reflects strong evidence of heterogeneity of response in both the cases and controls.

Alcoholism↗

PAWE-3D: visualizing power for association with error in case-control genetic studies of complex traits.

UNLABELLED: A website that plots power and sample size calculations over a range of up to eight parameters (including diagnostic misclassification error parameters) for two commonly used statistical tests of genetic association, the linear trend test and the genotypic test of association. AVAILABILITY: This method is made available via the website http://linkage.rockefeller.edu/pawe3d/ CONTACT: pawe3d@linkage.rockefeller.edu.

Algorithms↗

Power and sample size calculations in the presence of phenotype errors for case/control genetic association studies.

BACKGROUND: Phenotype error causes reduction in power to detect genetic association. We present a quantification of phenotype error, also known as diagnostic error, on power and sample size calculations for case-control genetic association studies between a marker locus and a disease phenotype. We consider the classic Pearson chi-square test for independence as our test of genetic association. To determine asymptotic power analytically, we compute the distribution's non-centrality parameter, which is a function of the case and control sample sizes, genotype frequencies, disease prevalence, and phenotype misclassification probabilities. We derive the non-centrality parameter in the presence of phenotype errors and equivalent formulas for misclassification cost (the percentage increase in minimum sample size needed to maintain constant asymptotic power at a fixed significance level for each percentage increase in a given misclassification parameter). We use a linear Taylor Series approximation for the cost of phenotype misclassification to determine lower bounds for the relative costs of misclassifying a true affected (respectively, unaffected) as a control (respectively, case). Power is verified by computer simulation. RESULTS: Our major findings are that: (i) the median absolute difference between analytic power with our method and simulation power was 0.001 and the absolute difference was no larger than 0.011; (ii) as the disease prevalence approaches 0, the cost of misclassifying a unaffected as a case becomes infinitely large while the cost of misclassifying an affected as a control approaches 0. CONCLUSION: Our work enables researchers to specifically quantify power loss and minimum sample size requirements in the presence of phenotype errors, thereby allowing for more realistic study design. For most diseases of current interest, verifying that cases are correctly classified is of paramount importance.

Alzheimer Disease↗

Time to remission and relapse after the first hospital admission in severe bipolar disorder.

BACKGROUND: Few studies of the time to remission and first relapse in severe bipolar disorder have been based on epidemiologically defined samples or have examined patient characteristics and time-varying indicators of medication use simultaneously. Using a cohort from the Suffolk County Mental Health Project, we describe these temporal patterns and their relationships with childhood, illness, and treatment characteristics. METHOD: A multi-facility cohort of 123 first-admission inpatients with DSM-IV bipolar disorder with psychotic features was followed for 4 years. Dates of the first complete remission (lasting at least 2 months), subsequent relapses, and use of antimanic (AM),antipsychotic (AP), and antidepressant (AD) medications were recorded. Childhood and illness characteristics were ascertained at baseline using standard instruments. RESULTS: By the 4-year point, 83.7% had achieved a full remission, with 42.3% remitting within 3 months, 63.4% within 6 months, and 74.8% within 1 year. Overall, younger age of onset, history of childhood psychopathology, and higher Brief Psychiatric Rating Scale (BPRS) anxiety/depression scores were significantly associated with longer time to remission. Discontinuing AM, AP and AD (compared to never using) and taking AP and AD (compared to never using) were significantly associated with remission in the multivariate analysis. Of the 103 participants with complete remission, 61.2% suffered a relapse; 24.3 % relapsed within 6 months of remission, and 35.9% within a year. Overall, 32.5% of the 123 participants had a single episode followed by full remission. Childhood internalizing-type problems, higher BPRS anxiety/depression and Hamilton depression scores, and an admission episode not involving mania, but not patterns of medication use, were associated with shorter time to relapse. CONCLUSION: By 4-year follow-up, the majority of severely ill bipolar patients had remitted from their initial episode, but more than half subsequently relapsed. Illness characteristics, especially depressive symptoms, and medication treatment were associated with the early course, although medication use after remission was not associated with relapse.

Adult↗

Factors affecting statistical power in the detection of genetic association.

The mapping of disease genes to specific loci has received a great deal of attention in the last decade, and many advances in therapeutics have resulted. Here we review family-based and population-based methods for association analysis. We define the factors that determine statistical power and show how study design and analysis should be designed to maximize the probability of localizing disease genes.

Data Interpretation, Statistical↗

Increasing power for tests of genetic association in the presence of phenotype and/or genotype error by use of double-sampling.

Phenotype and/or genotype misclassification can: significantly increase type II error probabilities for genetic case/control association, causing decrease in statistical power; and produce inaccurate estimates of population frequency parameters. We present a method, the likelihood ratio test allowing for errors (LRTae) that incorporates double-sample information for phenotypes and/or genotypes on a sub-sample of cases/controls. Population frequency parameters and misclassification probabilities are determined using a double-sample procedure as implemented in the Expectation-Maximization (EM) method. We perform null simulations assuming a SNP marker or a 4-allele (multi-allele) marker locus. To compare our method with the standard method that makes no adjustment for errors (LRTstd), we perform power simulations using a 2/k factorial design with high and low settings of: case/control samples, phenotype/genotype costs, double-sampled phenotypes/genotypes costs, phenotype/genotype error, and proportions of double-sampled individuals. All power simulations are performed fixing equal costs for the LRTstd and LRTae methods. We also consider case/control ApoE genotype data for an actual Alzheimer's study. The LRTae method maintains correct type I error proportions for all null simulations and all significance level thresholds (10%, 5%, 1%). LRTae average estimates of population frequencies and misclassification probabilities are equal to the true values, with variances of 10e-7 to 10e-8. For power simulations, the median power difference LRTae-LRTstd at the 5% significance level is 0.06 for multi-allele data and 0.01 for SNP data. For the ApoE data example, the LRTae and LRTstd p-values are 5.8 x 10e-5 and 1.6 x 10e-3, respectively. The increase in significance is due to adjustment in the LRTae for misclassification of the most commonly reported risk allele. We have developed freely available software that performs our LRTae statistic.

Journal Article↗

What SNP genotyping errors are most costly for genetic association studies?

Which genotype misclassification errors are most costly, in terms of increased sample size necessary (SSN) to maintain constant asymptotic power and significance level, when performing case/control studies of genetic association? We answer this question for single-nucleotide polymorphisms (SNPs), using the 2x3 chi(2) test of independence. Our strategy is to expand the noncentrality parameter of the asymptotic distribution of the chi(2) test under a specified alternative hypothesis to approximate SSN, using a linear Taylor series in the error parameters. We consider two scenarios: the first assumes Hardy-Weinberg equilibrium (HWE) for the true genotypes in both cases and controls, and the second assumes HWE only in controls. The Taylor series approximation has a relative error of less than 1% when each error rate is less than 2%. The most costly error is recording the more common homozygote as the less common homozygote, with indefinitely increasing cost coefficient as minor SNP allele frequencies approach 0 in both scenarios. The cost of misclassifying the more common homozygote to the heterozygote also becomes indefinitely large as the minor SNP allele frequency goes to 0 under both scenarios. For the violation of HWE modeled here, the cost of misclassifying a heterozygote to the less common homozygote becomes large, although bounded. Therefore, the use of SNPs with a small minor allele frequency requires careful attention to the frequency of genotyping errors to ensure that power specifications are met. Furthermore, the design of automated genotyping should minimize those errors whose cost coefficients can become indefinitely large.

Alleles↗

Percentiles of the null distribution of 2 maximum lod score tests.

We here consider the null distribution of the maximum lod score (LOD-M) obtained upon maximizing over transmission model parameters (penetrance values, dominance, and allele frequency) as well as the recombination fraction. Also considered is the lod score maximized over a fixed choice of genetic model parameters and recombination-fraction values set prior to the analysis (MMLS) as proposed by Hodge et al. The objective is to fit parametric distributions to MMLS and LOD-M. Our results are based on 3,600 simulations of samples of n = 100 nuclear families ascertained for having one affected member and at least one other sibling available for linkage analysis. Each null distribution is approximately a mixture p(2)(0) + (1 - p)(2)(v). The values of MMLS appear to fit the mixture 0.20(2)(0) + 0.80chi(2)(1.6). The mixture distribution 0.13(2)(0) + 0.87chi(2)(2.8). appears to describe the null distribution of LOD-M. From these results we derive a simple method for obtaining critical values of LOD-M and MMLS.

Alleles↗

Quantifying the percent increase in minimum sample size for SNP genotyping errors in genetic model-based association studies.

Kang et al. [Genet Epidemiol 2004;26:132-141] addressed the question of which genotype misclassification errors are most costly, in terms of minimum percentage increase in sample size necessary (%MSSN) to maintain constant asymptotic power and significance level, when performing case/control studies of genetic association in a genetic model-free setting. They answered the question for single nucleotide polymorphisms (SNPs) using the 2 x 3 chi2 test of independence. We address the same question here for a genetic model-based framework. The genetic model parameters considered are: disease model (dominant, recessive), genotypic relative risk, SNP (marker) and disease allele frequency, and linkage disequilibrium. %MSSN coefficients of each of the six possible error rates are determined by expanding the non-centrality parameter of the asymptotic distribution of the 2 x 3 chi2 test under a specified alternative hypothesis to approximate %MSSN using a linear Taylor series in the error rates. In this work we assume errors misclassifying one homozygote as another homozygote are 0, since these errors are thought to rarely occur in practice. Our findings are that there are settings of the genetic model parameters that lead to large total %MSSN for both dominant and recessive models. As SNP minor allele approaches 0, total %MSSN increases without bound, independent of other genetic model parameters. In general, %MSSN is a complex function of the genetic model parameters. Use of SNPs with small minor allele frequency requires careful attention to frequency of genotyping errors to insure that power specifications are met. Software to perform these calculations for study design is available, and an example of its use to study a disease is given.

Alleles↗

A survey of family care giving to elders in New York State: findings and implications.

It is estimated that there are 734,400 care giver households in New York State (9.6% [+/-] 0.8% of all households). Categorization of all care givers on a 5 level "intensity of care" measure reveals that, on average, care givers provide 22.1 hours of care per week. Highest intensity level 5 care givers (9.2% of all care givers), provide, on average, 88 hours of care per week and account for 36.3% of all care giving. The annualized market contribution of all care givers to the NYS health care system is estimated at between $7.5 and $11.2 billion dollars. The combination of care levels 4 and 5 contributed 70% of all care giving and account for about $5.2 billion in market value. Level 4 and 5 care givers are more likely than other care givers to report difficulties including the recent death of the care-receiver (p < 0.001), financial and employment setbacks (p < 0.001), emotional stress from worrying about the patient's future condition, dealing with cognitively impaired, physically unmanageable, opinionated, virtually immobile patient's, and insufficient support from family members (p < 0.05). Nevertheless, they also report such rewards as feeling grateful for improving the quality of the care receiver's life (43.5%) as well as love (17.2%). Of the nearly 15% of NYS care givers who used adult day care services, all reported that these services met their needs fully or partly. However, of the 85% who did not use the service, 33.5% were negative about it. And while use of adult day care services increases with intensity of care level, there is resistance to adult day care particularly among level 3 care givers, with negative statements from 46.8% of them.

Aged↗

Drug use and neurobehavioral, respiratory, and cognitive problems: precursors and mediators.

PURPOSE: To test a model of the early predictors and mediators of drug use and respiratory, neurobehavioral, and cognitive problems in adolescence and young adulthood. METHODS: We prospectively examined self-reported measures of unconventional behavior, peer- and self-drug use, and self-reported health problems in a sample of 286 males and 327 females. The sample represented the northeastern United States at the time the data were first collected in 1975. The participants were assessed in early, middle, and late adolescence and in young adulthood. Latent variable structural equation models were used to examine the data. RESULTS: Structural equation modeling conducted on the data provided support for the proposed longitudinal model. The findings indicated that adolescent drug use was associated indirectly with respiratory and directly with neurobehavioral and cognitive symptoms in young adulthood. Adolescent drug use during middle and late adolescence served as a mediator between unconventional behavior in early adolescence and health problems in young adulthood. CONCLUSIONS: A reduction in adolescent drug use may reduce respiratory and neurobehavioral and cognitive symptoms in young adulthood. This study identifies several points in the biopsychosocial pathways in adolescence leading to later health problems in young adulthood.

Adolescent↗