Adaptive Low-Rank Multi-Label Active Learning for Image Classification

Adaptive Low-Rank Multi-Label Active Learning for Image Classification
复制标题

DOI:
10.1145/3123266.3123388
复制
发表时间:
2017-10
期刊:
Proceedings of the 25th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Jian Wu;Anqian Guo;Victor S. Sheng;Pengpeng Zhao;Zhiming Cui;Hua Li
Jian Wu;Anqian Guo;Victor S. Sheng;Pengpeng Zhao;Zhiming Cui;Hua Li
中科院分区:
其他
文献类型:
--
作者:
Jian Wu;Anqian Guo;Victor S. Sheng;Pengpeng Zhao;Zhiming Cui;Hua Li

文献摘要

被引文献

相似文献

多标签主动学习在图像分类中的应用近年来引起了广泛的关注,相关的研究成果不断涌现。然而,现有的多标签主动学习算法没有反映样本数据的清洁性,标签相关性挖掘方法存在缺陷等问题。一方面,在实际应用中,样本数据往往是被污染的,这会干扰数据分布的估计,进而阻碍模型的训练。另一方面,以前的标签关系探索方法纯粹是基于观察到的标签分布的不完整的训练集,这不能提供足够有效的信息。为了解决这些问题,我们提出了一种新的自适应低秩多标签主动学习算法,称为LRMAL。具体来说,我们首先使用低秩矩阵恢复从噪声数据中学习有效的低秩特征表示。在随后的采样阶段,我们利用其优势,以评估每个未标记的示例标签对的一般信息。基于多标签数据集的样本空间与标签空间之间的内在映射关系,对训练集的不完整标签进行恢复,以实现更全面的标签相关性挖掘。此外,为了减少所选择的示例标签对之间的冗余,我们使用多样性测量多样化的采样数据。最后,一个有效的采样策略,通过整合这两个方面的潜在信息与不确定性的基础上自适应集成计划。实验结果证明了该方法的有效性。
Multi-label active learning for image classification has attracted great attention over recent years and a lot of relevant works are published continuously. However, there still remain some problems that need to be solved, such as existing multi-label active learning algorithms do not reflect on the cleanness of sample data and their ways on label correlation mining are defective. For one thing, sample data is usually contaminated in reality, which disturbs the estimation of data distribution and further hinders the model training. For another, previous approaches for label relationship exploration are purely based on the observed label distribution of an incomplete training set, which cannot provide sufficiently efficient information. To address these issues, we propose a novel adaptive low-rank multi-label active learning algorithm, called LRMAL. Specifically, we first use low-rank matrix recovery to learn an effective low-rank feature representation from the noisy data. In a subsequent sampling phase, we make use of its superiorities to evaluate the general informativeness of each unlabeled example-label pair. Based on an intrinsic mapping relation between the example space and the label space of a certain multi-label dataset, we recover the incomplete labels of a training set for a more comprehensive label correlation mining. Furthermore, to reduce the redundancy among the selected example-label pairs, we use a diversity measurement to diversify the sampled data. Finally, an effective sampling strategy is developed by integrating these two aspects of potential information with uncertainty based on an adaptive integration scheme. Experimental results demonstrate the effectiveness of our approach.