Hellinger distance decision trees are robust and skew-insensitive

Hellinger distance decision trees are robust and skew-insensitive
复制标题

DOI:
10.1007/s10618-011-0222-1
复制
发表时间:
2012-01-01
影响因子:
4.8
通讯作者:
Kegelmeyer, W. Philip
Kegelmeyer, W. Philip
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cieslak, David A.;Hoens, T. Ryan;Kegelmeyer, W. Philip

文献摘要

被引文献

相似文献

从不平衡数据中学习是一个重要而常见的问题。决策树,辅以抽样技术,已被证明是一种有效的方法来解决不平衡的数据问题。然而,尽管它们有效,抽样方法增加了复杂性和对参数选择的需要。为了绕过这些困难,我们提出了一种新的决策树技术,称为Hellinger距离决策树(HDDT),它使用Hellinger距离作为分裂标准。我们分析和经验证明了强大的斜不敏感性的Hellinger距离和它的优势,如熵(增益比)流行的替代品。我们应用了一个全面的经验评估框架,针对常用的采样和集成方法进行测试,考虑了58个不同数据集的性能。我们证明了HDDT在不平衡数据上的优越性(使用统计显著性的鲁棒性检验),以及它在平衡数据集上的竞争性能。因此,我们得出了特别实用的结论,对于不平衡的数据,使用Hellinger树与装袋(BG)没有任何抽样方法是足够的。我们在线提供了本文的所有数据集和软件(http:www.nd.edu/similar拨打/hddt)。
Learning from imbalanced data is an important and common problem. Decision trees, supplemented with sampling techniques, have proven to be an effective way to address the imbalanced data problem. Despite their effectiveness, however, sampling methods add complexity and the need for parameter selection. To bypass these difficulties we propose a new decision tree technique called Hellinger Distance Decision Trees (HDDT) which uses Hellinger distance as the splitting criterion. We analytically and empirically demonstrate the strong skew insensitivity of Hellinger distance and its advantages over popular alternatives such as entropy (gain ratio). We apply a comprehensive empirical evaluation framework testing against commonly used sampling and ensemble methods, considering performance across 58 varied datasets. We demonstrate the superiority (using robust tests of statistical significance) of HDDT on imbalanced data, as well as its competitive performance on balanced datasets. We thereby arrive at the particularly practical conclusion that for imbalanced data it is sufficient to use Hellinger trees with bagging (BG) without any sampling methods. We provide all the datasets and software for this paper online (http://www.nd.edu/similar to dial/hddt).