Predictor correlation impacts machine learning algorithms: implications for genomic studies

Predictor correlation impacts machine learning algorithms: implications for genomic studies
复制标题

DOI:
10.1093/bioinformatics/btp331
复制
发表时间:
2009-08-01
期刊:
影响因子:
5.8
通讯作者:
Malley, James D.
Malley, James D.
中科院分区:
生物学3区
文献类型:
--
作者:
Nicodemus, Kristin K.;Malley, James D.

文献摘要

被引文献

相似文献

动机:高通量基因组学的出现产生了具有大量预测因素的研究(例如全基因组关联、微阵列研究)。机器学习算法(MLA)是在高维数据中识别表型相关变量的一种高效的计算方法。有来自数学理论的重要结果和大量的实践结果证明了它们的价值。MLA的一个吸引人的特点是,许多MLA在完全多变量的环境中运行,允许在它们协同行动时包括较小的重要变量。然而,基因组相关数据中常见条件下MLA的某些性质还没有得到很好的研究,尤其是预测者之间的相关性带来了一个问题。结果:通过广泛的模拟,我们证明了考虑预测者之间的相关性对于使用三个MLA的变量重要性度量(VIMS)进行有效推断至关重要:随机林(RF)、条件推理林(CIF)和蒙特卡罗逻辑回归(MCLR)。使用病例对照图示,我们表明,在复杂疾病研究中遇到的效应大小时,RF VIMS-即使基于排列-比其他算法检测关联的能力更差。当“因果”预测因素与其他预测因素相关时,这种减少就会发生,当使用基尼指数构建射频树时,这种减少的幅度最大。事实上,当树的终端节点较小时,RF GINI VIM在相关性下是有偏差的,依赖于预测器相关性强度/数量,并且过度训练以适应数据的随机波动。基于排列的VIM分布对于相关的预测因素来说变量较少,并且是无偏的,因此当预测因素相关时可能是首选的。MLA是高维数据分析的强大工具,但必须经过深思熟虑地使用算法才能得出有效的结论。
Motivation: The advent of high-throughput genomics has produced studies with large numbers of predictors (e. g. genome-wide association, microarray studies). Machine learning algorithms (MLAs) are a computationally efficient way to identify phenotype-associated variables in high-dimensional data. There are important results from mathematical theory and numerous practical results documenting their value. One attractive feature of MLAs is that many operate in a fully multivariate environment, allowing for small-importance variables to be included when they act cooperatively. However, certain properties of MLAs under conditions common in genomic-related data have not been well-studied-in particular, correlations among predictors pose a problem.Results: Using extensive simulation, we showed considering correlation within predictors is crucial in making valid inferences using variable importance measures (VIMs) from three MLAs: random forest (RF), conditional inference forest (CIF) and Monte Carlo logic regression (MCLR). Using a case-control illustration, we showed that the RF VIMs-even permutation-based-were less able to detect association than other algorithms at effect sizes encountered in complex disease studies. This reduction occurred when 'causal' predictors were correlated with other predictors, and was sharpest when RF tree building used the Gini index. Indeed, RF Gini VIMs are biased under correlation, dependent on predictor correlation strength/number and over-trained to random fluctuations in data when tree terminal node size was small. Permutation-based VIM distributions were less variable for correlated predictors and are unbiased, thus may be preferred when predictors are correlated. MLAs are a powerful tool for high-dimensional data analysis, but well-considered use of algorithms is necessary to draw valid conclusions.