Variable Importance Using Decision Trees

Variable Importance Using Decision Trees
复制标题

使用决策树的变量重要性

DOI:
--
复制
发表时间:
2017
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Ameet Talwalkar
Ameet Talwalkar
中科院分区:
--
文献类型:
--
作者:
S. J. Kazemitabar;A. Amini;Adam Bloniarz;Ameet Talwalkar

文献摘要

被引文献

相似文献

决策树和随机森林是公认的模型,它们不仅提供了良好的预测性能,而且提供了丰富的特征重要性信息。虽然实践者经常使用依赖于这种基于杂质的信息的可变重要性方法,但从理论角度来看,这些方法的特点仍然很差。我们对这些方法的性能提供了新的见解,通过在各种建模假设下推导出高维环境下的有限样本性能保证。我们通过大量的模拟进一步证明了这些基于杂质的方法的有效性。
Decision trees and random forests are well established models that not only offer good predictive performance, but also provide rich feature importance information. While practitioners often employ variable importance methods that rely on this impurity-based information, these methods remain poorly characterized from a theoretical perspective. We provide novel insights into the performance of these methods by deriving finite sample performance guarantees in a high-dimensional setting under various modeling assumptions. We further demonstrate the effectiveness of these impurity-based methods via an extensive set of simulations.