A random forests quantile classifier for class imbalanced data

A random forests quantile classifier for class imbalanced data
复制标题

DOI:
10.1016/j.patcog.2019.01.036
复制
发表时间:
2019-06-01
影响因子:
8
通讯作者:
Ishwaran, Hemant
Ishwaran, Hemant
中科院分区:
计算机科学1区
文献类型:
--
作者:
O'Brien, Robert;Ishwaran, Hemant

文献摘要

被引文献

相似文献

扩展以前的工作分位数分类器(q-分类器),我们提出了类不平衡问题的q*-分类器。如果少数类条件概率超过0 < q* < 1,则分类器将样本分配给少数类,其中g* 等于观察少数类样本的无条件概率。q*-分类的动机源于基于密度的方法,并导致q*-分类器最大化真阳性率和真阴性率之和的有用属性。此外,由于该过程可以等效地表示为成本加权贝叶斯分类器,因此它也使加权风险最小化。由于这种双重优化,q* 分类器可以在不平衡问题中实现接近零的风险,同时优化真阳性率和真阴性率。我们使用随机森林来应用q*-分类。这种新的方法,我们称之为G-means的性能和变量选择优于现有的技术或具有竞争力。也被认为是多类不平衡设置的扩展。(C)2019爱思唯尔有限公司版权所有。
Extending previous work on quantile classifiers (q-classifiers) we propose the q*-classifier for the class imbalance problem. The classifier assigns a sample to the minority class if the minority class conditional probability exceeds 0 < q* < 1, where g* equals the unconditional probability of observing a minority class sample. The motivation for q*-classification stems from a density-based approach and leads to the useful property that the q*-classifier maximizes the sum of the true positive and true negative rates. Moreover, because the procedure can be equivalently expressed as a cost-weighted Bayes classifier, it also minimizes weighted risk. Because of this dual optimization, the q*-classifier can achieve near zero risk in imbalance problems, while simultaneously optimizing true positive and true negative rates. We use random forests to apply q*-classification. This new method which we call RFQ is shown to outperform or is competitive with existing techniques with respect to G-mean performance and variable selection. Extensions to the multiclass imbalanced setting are also considered. (C) 2019 Elsevier Ltd. All rights reserved.