Improving imbalanced classification using near-miss instances

Improving imbalanced classification using near-miss instances
复制标题

DOI:
10.1016/j.eswa.2022.117130
复制
发表时间:
2022-04
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Akira Tanimoto;S. Yamada;Takashi Takenouchi;Masashi Sugiyama;H. Kashima
Akira Tanimoto;S. Yamada;Takashi Takenouchi;Masashi Sugiyama;H. Kashima
中科院分区:
其他
文献类型:
--
作者:
Akira Tanimoto;S. Yamada;Takashi Takenouchi;Masashi Sugiyama;H. Kashima

文献摘要

相似文献

类不平衡是分类中的一个主要问题,即罕见类(正)的样本量通常是性能瓶颈。然而,在现实世界的情况下,“差一点”的积极实例,即消极但近乎积极的实例,有时很多。例如,洪水等自然灾害很少发生,而相对较多的险情是,虽然没有发生真正的洪水,但水位接近河岸高度。我们表明,即使真正的正面案例非常有限,例如在灾害预测中,也可以通过获得精细的类似标签的侧面信息“正面”(例如,河流的水位)来提高准确性,从而将侥幸案例与其他负面案例区分开来。传统的代价敏感分类不能利用这种侧信息,而且正样本的小样本导致估计方差大。我们的方法与使用特权信息(LUPI)的学习是一致的,它利用侧信息进行训练,而不预测侧信息本身。我们从理论上证明,我们的方法减少了估计方差,提供了大量的近射正实例,以换取额外的偏差。广泛的实验结果表明,我们的方法往往优于或比较有利的现有方法。
The class imbalance is a major issue in classification, i.e., the sample size of a rare class (positive) is often a performance bottleneck. In real-world situations, however, “near-miss” positive instances, i.e., negative but nearly-positive instances, are sometimes plentiful. For example, natural disasters such as floods are rare, while there are relatively plentiful near-miss cases where actual floods did not occur but the water level approached the bank height. We show that even when the true positive cases are quite limited, such as in disaster forecasting, the accuracy can be improved by obtaining refined label-like side-information “positivity” (e.g., the water level of the river) to distinguish near-miss cases from other negatives. Conventional cost-sensitive classification cannot utilize such side-information, and the small size of the positive sample causes high estimation variance. Our approach is in line with learning using privileged information (LUPI), which exploits side-information for training without predicting the side-information itself. We theoretically prove that our method reduces the estimation variance, provided that near-miss positive instances are plentiful, in exchange for additional bias. Results of extensive experiments demonstrate that our method tends to outperform or compares favorably to existing approaches.