Permutation importance: a corrected feature importance measure

Permutation importance: a corrected feature importance measure
复制标题

DOI:
10.1093/bioinformatics/btq134
复制
发表时间:
2010-05-15
期刊:
影响因子:
5.8
通讯作者:
Lengauer, Thomas
Lengauer, Thomas
中科院分区:
生物学3区
文献类型:
--
作者:
Altmann, Andre;Tolosi, Laura;Lengauer, Thomas

文献摘要

被引文献

相似文献

动机:在生命科学中,机器学习模型的可解释性与其预测准确性一样重要。线性模型可能是最常用的方法来评估功能的相关性,尽管他们的相对可扩展性。然而,在过去的几年中,有效的特征相关性的估计已经得到了高度复杂的或非参数模型,如支持向量机和随机森林(RF)模型。最近,它已被观察到,RF模型是有偏见的,在这样一种方式,具有大量的categories are preferred.Results:在这项工作中,我们引入了一个启发式规范化功能的重要性措施,可以纠正功能的重要性偏差。该方法基于结果向量的重复排列,用于估计非信息设置中每个变量的测量重要性的分布。观察到的重要性的P值提供了特征重要性的校正度量。我们将我们的方法应用于模拟数据,并证明:(i)非信息性预测因子不会获得显著的P值,(ii)信息性变量可以成功地在非信息性变量中恢复,(iii)用排列重要性(PIMP)计算的P值非常有助于确定变量的显著性,从而提高模型的可解释性。此外,PIMP被用来纠正RF为基础的重要性措施,为两个现实世界的案例研究。我们提出了一个改进的RF模型,使用的显着变量的PIMP措施,并表明其预测精度是上级的其他现有models.Availability:R代码的方法在这篇文章中提供的是在http://www.mpi-inf.mpg.de/similar到altmann/download/PIMP。RContact:altmann@mpi-inf.mpg.de,laura.tolosi@mpi-inf.mpg. de补充信息:补充数据可在生物信息学在线。
Motivation: In life sciences, interpretability of machine learning models is as important as their prediction accuracy. Linear models are probably the most frequently used methods for assessing feature relevance, despite their relative inflexibility. However, in the past years effective estimators of feature relevance have been derived for highly complex or non-parametric models such as support vector machines and RandomForest (RF) models. Recently, it has been observed that RF models are biased in such a way that categorical variables with a large number of categories are preferred.Results: In this work, we introduce a heuristic for normalizing feature importance measures that can correct the feature importance bias. The method is based on repeated permutations of the outcome vector for estimating the distribution of measured importance for each variable in a non-informative setting. The P-value of the observed importance provides a corrected measure of feature importance. We apply our method to simulated data and demonstrate that (i) non-informative predictors do not receive significant P-values, (ii) informative variables can successfully be recovered among non-informative variables and (iii) P-values computed with permutation importance (PIMP) are very helpful for deciding the significance of variables, and therefore improve model interpretability. Furthermore, PIMP was used to correct RF-based importance measures for two real-world case studies. We propose an improved RF model that uses the significant variables with respect to the PIMP measure and show that its prediction accuracy is superior to that of other existing models.Availability: R code for the method presented in this article is available at http://www.mpi-inf.mpg.de/similar to altmann/download/PIMP.RContact: altmann@mpi-inf.mpg.de, laura.tolosi@mpi-inf.mpg.deSupplementary information: Supplementary data are available at Bioinformatics online.