Screening large-scale association study data: exploiting interactions using random forests.

Screening large-scale association study data: exploiting interactions using random forests.
复制标题

DOI:
10.1186/1471-2156-5-32
复制
发表时间:
2004-12-10
期刊:
影响因子:
2.9
通讯作者:
Van Eerdewegh P
Van Eerdewegh P
中科院分区:
生物学3区
文献类型:
--
作者:
Lunetta KL;Hayward LB;Segal J;Van Eerdewegh P

文献摘要

参考文献

被引文献

相似文献

针对复杂疾病的全基因组关联研究将产生数十万个单核苷酸多态性 (SNP) 的基因型。处理大量 SNP 的第一个合乎逻辑的方法是使用一些测试来筛选 SNP,仅保留那些符合某些标准的 SNP 以供进一步研究。例如,SNP 可以按 p 值排序,并保留 p 值最低的那些。当 SNP 在群体中具有较大的相互作用效应但较小的边际效应时,当使用单变量检验进行筛选时,它们不太可能被保留。然而,对于具有数千个 SNP 的数据集来说,预先指定相互作用的基于模型的屏幕是不切实际的。随机森林分析是一种替代方法,它为每个预测变量生成单一的重要性度量,该方法考虑变量之间的相互作用,而不需要模型规范。交互作用增加了各个交互变量的重要性,使它们相对于其他变量更有可能被赋予高度重要性。我们使用具有多达 32 个基因座的复杂疾病模型,结合遗传异质性和多基因座相互作用,测试随机森林作为筛选程序的性能,以从大量不相关的 SNP 中识别少量风险相关的 SNP。在保持其他因素不变的情况下,如果风险 SNP 相互作用,则随机森林重要性度量作为筛选工具的性能显着优于 Fisher Exact 检验。随着相互作用的 SNP 数量的增加,随机森林分析相对于 Fisher Exact 筛选检验的性能改进也随之增加。当分析中的 SNP 不相互作用时,随机森林的表现类似于单变量 Fisher Exact 检验作为筛选工具。在大规模遗传关联研究中,真正的风险相关 SNP 或 SNP 与环境协变量之间存在未知的相互作用,与标准单变量筛选方法相比,使用随机森林分析筛选 SNP 可以显着减少需要保留用于进一步研究的 SNP 数量。
Genome-wide association studies for complex diseases will produce genotypes on hundreds of thousands of single nucleotide polymorphisms (SNPs). A logical first approach to dealing with massive numbers of SNPs is to use some test to screen the SNPs, retaining only those that meet some criterion for futher study. For example, SNPs can be ranked by p-value, and those with the lowest p-values retained. When SNPs have large interaction effects but small marginal effects in a population, they are unlikely to be retained when univariate tests are used for screening. However, model-based screens that pre-specify interactions are impractical for data sets with thousands of SNPs. Random forest analysis is an alternative method that produces a single measure of importance for each predictor variable that takes into account interactions among variables without requiring model specification. Interactions increase the importance for the individual interacting variables, making them more likely to be given high importance relative to other variables. We test the performance of random forests as a screening procedure to identify small numbers of risk-associated SNPs from among large numbers of unassociated SNPs using complex disease models with up to 32 loci, incorporating both genetic heterogeneity and multi-locus interaction. Keeping other factors constant, if risk SNPs interact, the random forest importance measure significantly outperforms the Fisher Exact test as a screening tool. As the number of interacting SNPs increases, the improvement in performance of random forest analysis relative to Fisher Exact test for screening also increases. Random forests perform similarly to the univariate Fisher Exact test as a screening tool when SNPs in the analysis do not interact. In the context of large-scale genetic association studies where unknown interactions exist among true risk-associated SNPs or SNPs and environmental covariates, screening SNPs using random forest analyses can significantly reduce the number of SNPs that need to be retained for further study compared to standard univariate screening methods.
DOI: 10.1002/gepi.2001.21.s1.s626
发表时间: 2001-01-01
影响因子: 2.1
作者:
Kooperberg, C;Ruczinski, I;Hsu, L
通讯作者: Hsu, L
DOI: 10.1086/321276
发表时间: 2001-07-01
影响因子: 9.8
作者:
Ritchie, MD;Hahn, LW;Moore, JH
通讯作者: Moore, JH
DOI: 10.1016/s0065-2660(01)42028-1
发表时间: 2001-01-01
期刊: GENETIC DISSECTION OF COMPLEX TRAITS
影响因子: --
作者:
Province, MA;Shannon, WD;Rao, DC
通讯作者: Rao, DC
DOI: 10.1002/gepi.2001.21.s1.s649
发表时间: 2001-01-01
影响因子: 2.1
作者:
York, TP;Eaves, LJ
通讯作者: Eaves, LJ
使用随机森林绘制复杂性状。
DOI: 10.1186/1471-2156-4-s1-s64
发表时间: 2003-12-31
期刊: BMC GENETICS
影响因子: 2.9
作者:
Bureau, A;Dupuis, J;Hayward, B;Falls, K;Van Eerdewegh, P
通讯作者: Van Eerdewegh, P