SNP interaction detection with Random Forests in high-dimensional genetic data.

SNP interaction detection with Random Forests in high-dimensional genetic data.
复制标题

DOI:
10.1186/1471-2105-13-164
复制
发表时间:
2012-07-15
期刊:
影响因子:
3
通讯作者:
Biernacka JM
Biernacka JM
中科院分区:
生物学4区
文献类型:
--
作者:
Winham SJ;Colby CL;Freimuth RR;Wang X;de Andrade M;Huebner M;Biernacka JM

文献摘要

参考文献

被引文献

相似文献

在高维数据中识别与复杂人类特征相关的变异是全基因组关联研究的中心目标。然而,在这些研究中,单变量分析通常忽略了复杂的病因学,如基因-基因相互作用。随机森林(RF)是一种流行的数据挖掘技术,可以容纳大量的预测变量,并允许复杂的模型相互作用。RF分析产生变量重要性的度量,可用于对预测变量进行排序。因此,使用RF的单核苷酸多态性(SNP)分析作为一种考虑高维数据中相互作用的潜在过滤方法越来越受欢迎。然而,数据维度对RF识别相互作用的能力的影响还没有得到彻底的探讨。我们调查的能力,排名从变量的重要性措施,以检测基因-基因相互作用的影响和他们的潜在有效性相比,从单变量逻辑回归的p值过滤器,特别是随着数据变得越来越高维。RF有效地识别低维数据中的相互作用。随着预测变量总数的增加,相互作用SNP的检测概率比非相互作用SNP下降得更快,这表明在高维数据中,RF变量重要性度量捕获的是边际效应,而不是捕获相互作用的效应。虽然RF仍然是一个很有前途的数据挖掘技术,扩展单变量方法的条件下,同时对多个变量,RF变量的重要性措施未能检测到在高维数据中的相互作用的影响,在没有一个强大的边缘组件,因此可能不是有用的过滤器技术,允许在全基因组数据的相互作用的影响。
Identifying variants associated with complex human traits in high-dimensional data is a central goal of genome-wide association studies. However, complicated etiologies such as gene-gene interactions are ignored by the univariate analysis usually applied in these studies. Random Forests (RF) are a popular data-mining technique that can accommodate a large number of predictor variables and allow for complex models with interactions. RF analysis produces measures of variable importance that can be used to rank the predictor variables. Thus, single nucleotide polymorphism (SNP) analysis using RFs is gaining popularity as a potential filter approach that considers interactions in high-dimensional data. However, the impact of data dimensionality on the power of RF to identify interactions has not been thoroughly explored. We investigate the ability of rankings from variable importance measures to detect gene-gene interaction effects and their potential effectiveness as filters compared to p-values from univariate logistic regression, particularly as the data becomes increasingly high-dimensional. RF effectively identifies interactions in low dimensional data. As the total number of predictor variables increases, probability of detection declines more rapidly for interacting SNPs than for non-interacting SNPs, indicating that in high-dimensional data the RF variable importance measures are capturing marginal effects rather than capturing the effects of interactions. While RF remains a promising data-mining technique that extends univariate methods to condition on multiple variables simultaneously, RF variable importance measures fail to detect interaction effects in high-dimensional data in the absence of a strong marginal component, and therefore may not be useful as a filter technique that allows for interaction effects in genome-wide data.
DOI: 10.1038/nrg2579
发表时间: 2009-06
期刊: Nature reviews. Genetics
影响因子: --
作者:
Cordell HJ
通讯作者: Cordell HJ
DOI: 10.1086/338759
发表时间: 2002-02-01
影响因子: 9.8
作者:
Culverhouse, R;Suarez, BK;Reich, T
通讯作者: Reich, T
DOI: 10.1086/321276
发表时间: 2001-07-01
影响因子: 9.8
作者:
Ritchie, MD;Hahn, LW;Moore, JH
通讯作者: Moore, JH
DOI: 10.1093/bioinformatics/bti689
发表时间: 2005-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Montana, G
通讯作者: Montana, G
DOI: 10.1093/bioinformatics/btp331
发表时间: 2009-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Nicodemus, Kristin K.;Malley, James D.
通讯作者: Malley, James D.