PubMed Health⌕ Search

PubMed · 15588316

Screening large-scale association study data: exploiting interactions using random forests.

Abstract

BACKGROUND: Genome-wide association studies for complex diseases will produce genotypes on hundreds of thousands of single nucleotide polymorphisms (SNPs). A logical first approach to dealing with massive numbers of SNPs is to use some test to screen the SNPs, retaining only those that meet some criterion for further study. For example, SNPs can be ranked by p-value, and those with the lowest p-values retained. When SNPs have large interaction effects but small marginal effects in a population, they are unlikely to be retained when univariate tests are used for screening. However, model-based screens that pre-specify interactions are impractical for data sets with thousands of SNPs. Random forest analysis is an alternative method that produces a single measure of importance for each predictor variable that takes into account interactions among variables without requiring model specification. Interactions increase the importance for the individual interacting variables, making them more likely to be given high importance relative to other variables. We test the performance of random forests as a screening procedure to identify small numbers of risk-associated SNPs from among large numbers of unassociated SNPs using complex disease models with up to 32 loci, incorporating both genetic heterogeneity and multi-locus interaction. RESULTS: Keeping other factors constant, if risk SNPs interact, the random forest importance measure significantly outperforms the Fisher Exact test as a screening tool. As the number of interacting SNPs increases, the improvement in performance of random forest analysis relative to Fisher Exact test for screening also increases. Random forests perform similarly to the univariate Fisher Exact test as a screening tool when SNPs in the analysis do not interact. CONCLUSIONS: In the context of large-scale genetic association studies where unknown interactions exist among true risk-associated SNPs or SNPs and environmental covariates, screening SNPs using random forest analyses can significantly reduce the number of SNPs that need to be retained for further study compared to standard univariate screening methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kathryn L Lunetta, L Brooke Hayward, Jonathan Segal, Paul Van Eerdewegh. 2004-12-10. Screening large-scale association study data: exploiting interactions using random forests.. https://doi.org/10.1186/1471-2156-5-32

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Comprehensive evaluation of genetic variation in the IGF1 gene and risk of prostate cancer.

Insulin-like growth factor-I (IGF1) stimulates cell proliferation, decreases apoptosis, and has been implicated in cancer development. Epidemiological studies have shown elevated levels of circulating IGF1 to be associated with increased risk of prostate cancer. To what extent genetic variation in the IGF1 gene is related to prostate cancer risk is largely unknown. We performed a comprehensive haplotype tagging (HT) assessment of single nucleotide polymorphisms (SNPs) representing the common haplotype variation in the IGF1 gene. We genotyped 10 SNPs (9 haplotype tagging SNPs (htSNPs)) within Cancer Prostate in Sweden (CAPS), a case-control study of 2,863 cases and 1,737 controls, in order to investigate if genetic variation in the IGF1 gene is associated with prostate cancer risk. Three haplotype blocks were identified across the IGF1 gene and 9 SNPs were selected as haplotype tagging SNPs. Common haplotypes in the block covering the 3' region of the IGF1 gene showed significant global association with prostate cancer risk (p = 0.004), with one particular haplotype giving an odds ratio of 1.46 (95% CI = 1.15-1.84, p = 0.002). This haplotype had a prevalence of 5% in the study population. Our results indicate that common variation in the IGF1 gene, particularly in the 3' region, may affect prostate cancer risk. Further studies on genetic variations in the IGF1 gene in relation to prostate cancer risk as well as to circulating levels of IGF1 are needed to confirm this novel finding.

Case-Control Studies↗

Exploring the joint effects of silicosis and smoking on lung cancer risks.

Cigarette smoking and silicosis are potential causes of lung cancer among workers exposed to silica dust, but their joint effects are unclear. We explored the possible interactions between silicosis and smoking on lung cancer risks by summarizing data from the published literature. The standardized mortality ratio or standardized incidence ratio reported in each published report was first adjusted using "smoking adjustment factors" to correct for the biased estimation of the expected numbers of lung cancer among smokers and nonsmokers when using general population rates in the indirect standardization process. The ratio of the effect of silicosis on lung cancer risk among smokers to that among nonsmoker was calculated and named the "relative silicosis effect (RSE)". The synergy index was estimated to assess the additive interaction. Metaanalyses were used to obtain the weighed means of the RSE and synergy index. Ten cohort studies were reviewed and combined to yield a weighed RSE of 0.29 (95% CI: 0.20, 0.42), indicating negative risk-ratio multiplication between smoking and silicosis on the lung cancer risk. The combined weighed synergy index was 1.00 (95% CI: 0.79, 1.26), suggesting no departure from additivity. Sensitivity analyses showed that both estimates were quite robust. The independent risk-ratio effect of silicosis on lung cancer in smokers was about 30% of that in nonsmokers, and the joint effects of smoking and silicosis on the risk of lung cancer did not deviate from additivity and hence did not support biological synergism/antagonism.

Case-Control Studies↗