PubMed Health⌕ Search

PubMed · 16077740

Adjusting multiple testing in multilocus analyses using the eigenvalues of a correlation matrix.

Abstract

Correlated multiple testing is widely performed in genetic research, particularly in multilocus analyses of complex diseases. Failure to control appropriately for the effect of multiple testing will either result in a flood of false-positive claims or in true hits being overlooked. Cheverud proposed the idea of adjusting correlated tests as if they were independent, according to an 'effective number' (M(eff)) of independent tests. However, our experience has indicated that Cheverud's estimate of the Meff is overly large and will lead to excessively conservative results. We propose a more accurate estimate of the M(eff), and design M(eff)-based procedures to control the experiment-wise significant level and the false discovery rate. In an evaluation, based on both real and simulated data, the M(eff)-based procedures were able to control the error rate accurately and consequently resulted in a power increase, especially in multilocus analyses. The results confirm that the M(eff) is a useful concept in the error-rate control of correlated tests. With its efficiency and accuracy, the M(eff) method provides an alternative to computationally intensive methods such as the permutation test.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J Li, L Ji. 2005. Adjusting multiple testing in multilocus analyses using the eigenvalues of a correlation matrix.. https://doi.org/10.1038/sj.hdy.6800717

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Non-normal path analysis in the presence of measurement error and missing data: a Bayesian analysis of nursing homes' structure and outcomes.

Path analytic models are useful tools in quantitative nursing research. They allow researchers to hypothesize causal inferential paths and test the significance of these paths both directly and indirectly through a mediating variable. A standard statistical method in the path analysis literature is to treat the variables as having a normal distribution and to estimate paths using several least squares regression equations. The parameters corresponding to the direct paths have point and interval estimates based on normal distribution theory. Indirect paths are a product of the direct path from the independent variable to the mediating variable and the direct path of the mediating variable to the dependent variable. However, in the case of non-normal distributions, the point and interval estimates of the indirect path become much more difficult to estimate. We address the issue of calculating indirect point and interval estimates in the case of non-normally distributed data. Our substantive application is a nursing home research problem in which the variables in the path analysis of interest involve variables with normal, Bernoulli, or Poisson distributions. Additionally, one of the Poisson variables is observed with error. This paper addresses estimating point and interval estimation of indirect paths for variables with non-normal distributions in the presence of missing data and measurement error. We handle these difficulties from a fully Bayesian point of view. We present our substantive path analysis motivated from a nursing home structure, process, and outcomes model. Our results focus on the impact job turnover in the nursing homes has on nursing home outcomes.

Data Interpretation, Statistical↗

Improving ecological inference using individual-level data.

In typical small-area studies of health and environment we wish to make inference on the relationship between individual-level quantities using aggregate, or ecological, data. Such ecological inference is often subject to bias and imprecision, due to the lack of individual-level information in the data. Conversely, individual-level survey data often have insufficient power to study small-area variations in health. Such problems can be reduced by supplementing the aggregate-level data with small samples of data from individuals within the areas, which directly link exposures and outcomes. We outline a hierarchical model framework for estimating individual-level associations using a combination of aggregate and individual data. We perform a comprehensive simulation study, under a variety of realistic conditions, to determine when aggregate data are sufficient for accurate inference, and when we also require individual-level information. Finally, we illustrate the methods in a case study investigating the relationship between limiting long-term illness, ethnicity and income in London.

Data Interpretation, Statistical↗

Modelling SARS data using threshold geometric process.

During the outbreak of an epidemic disease, for example, the severe acute respiratory syndrome (SARS), the number of daily infected cases often exhibit multiple trends: monotone increasing during the growing stage, stationary during the stabilized stage and then decreasing during the declining stage. Lam first proposed modelling a monotone trend by a geometric process (GP) [X(i), i=1,2,...] directly such that [a(i-1)X(i), i=1,2,...] forms a renewal process for some ratio a>0 which measures the direction and strength of the trend. Parameters can be conveniently estimated using the LSE methods. Previous GP models limit to data with only a single trend. For data with multiple trends, we propose a moving window technique to locate the turning point(s). The threshold GP model is fitted to the SARS data from four regions in 2003.

Data Interpretation, Statistical↗