Transfer Learning for Class Imbalance Problems with Inadequate Data.

Transfer Learning for Class Imbalance Problems with Inadequate Data.
复制标题

DOI:
10.1007/s10115-015-0870-3
复制
发表时间:
2016-07
影响因子:
2.7
通讯作者:
Reddy CK
Reddy CK
中科院分区:
计算机科学4区
文献类型:
--
作者:
Al-Stouhi S;Reddy CK

文献摘要

被引文献

相似文献

数据挖掘中的一个基本问题是在数据分布不均匀的情况下有效地建立鲁棒的分类器。类不平衡分类器是专门针对偏态分布数据集训练的。现有的方法假设大量的训练样本作为构建有效分类器的基本先决条件。然而,当不容易获得足够的数据时,由于类之间的不平等分布,代表性分类算法的开发变得更加困难。我们提供了一个统一的框架,该框架将使用迁移学习机制来利用辅助数据,同时构建一个强大的分类器,以解决在特定目标领域中存在少量训练样本的情况下的不平衡问题。迁移学习方法在训练样本不足时使用辅助数据来增强学习,在本文中,我们将开发一种方法,该方法经过优化,可以同时增强训练数据并在偏斜数据集中引入平衡。我们提出了一种新的基于提升的实例转移分类器,该分类器具有依赖于标签的更新机制,该机制同时补偿类不平衡,并结合来自辅助域的样本来改善分类。我们提供了我们的方法的理论和经验验证,并适用于医疗保健和文本分类应用。
A fundamental problem in data mining is to effectively build robust classifiers in the presence of skewed data distributions. Class imbalance classifiers are trained specifically for skewed distribution datasets. Existing methods assume an ample supply of training examples as a fundamental prerequisite for constructing an effective classifier. However, when sufficient data is not readily available, the development of a representative classification algorithm becomes even more difficult due to the unequal distribution between classes. We provide a unified framework that will potentially take advantage of auxiliary data using a transfer learning mechanism and simultaneously build a robust classifier to tackle this imbalance issue in the presence of few training samples in a particular target domain of interest. Transfer learning methods use auxiliary data to augment learning when training examples are not sufficient and in this paper we will develop a method that is optimized to simultaneously augment the training data and induce balance into skewed datasets. We propose a novel boosting based instance-transfer classifier with a label-dependent update mechanism that simultaneously compensates for class imbalance and incorporates samples from an auxiliary domain to improve classification. We provide theoretical and empirical validation of our method and apply to healthcare and text classification applications.