Interpretable Minority Synthesis for Imbalanced Classification

Interpretable Minority Synthesis for Imbalanced Classification
复制标题

DOI:
10.24963/ijcai.2021/350
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Yi He;Fudong Lin;Xu Yuan;N. Tzeng
Yi He;Fudong Lin;Xu Yuan;N. Tzeng
中科院分区:
其他
文献类型:
--
作者:
Yi He;Fudong Lin;Xu Yuan;N. Tzeng

文献摘要

被引文献

相似文献

本文提出了一种新的过采样方法,力求平衡类先验与相当不平衡的高维数据分布。我们的方法的关键在于学习可解释的潜在表示,可以通过使用生成对抗网络(GAN)来模拟少数样本的合成机制。一个贝叶斯正则化器被用来引导GAN提取一组突出的特征,这些特征要么是解开的,要么是密集纠缠的,它们的相互作用由一个规定的结构控制,由人在回路中定义。因此,我们的GAN具有更高的样本复杂度,即使在训练期间少数类的大小非常小,也能够合成高质量的少数样本。实证研究证实,我们的方法可以使简单的分类器,以实现上级的不平衡分类性能超过国家的最先进的竞争对手,是强大的各种不平衡设置。代码发布在github.com/fudonglin/IMSIC。
This paper proposes a novel oversampling approach that strives to balance the class priors with a considerably imbalanced data distribution of high dimensionality. The crux of our approach lies in learning interpretable latent representations that can model the synthetic mechanism of the minority samples by using a generative adversarial network(GAN). A Bayesian regularizer is imposed to guide the GAN to extract a set of salient features that are either disentangled or intensionally entangled, with their interplay controlled by a prescribed structure, defined with human-in-the-loop. As such, our GAN enjoys an improved sample complexity, being able to synthesize high-quality minority samples even if the sizes of minority classes are extremely small during training. Empirical studies substantiate that our approach can empower simple classifiers to achieve superior imbalanced classification performance over the state-of-the-art competitors and is robust across various imbalance settings. Code is released in github.com/fudonglin/IMSIC.