Combining active learning and semi-supervised learning to construct SVM classifier

Combining active learning and semi-supervised learning to construct SVM classifier
复制标题

DOI:
10.1016/j.knosys.2013.01.032
复制
发表时间:
2013-05-01
影响因子:
8.8
通讯作者:
Qi, Guanghui
Qi, Guanghui
中科院分区:
计算机科学1区
文献类型:
--
作者:
Leng, Yan;Xu, Xinyan;Qi, Guanghui

文献摘要

被引文献

相似文献

大多数分类算法的一个关键问题是它们需要大量的标记样本来训练分类器。由于人工标注耗时较长,研究人员提出了主动学习和半监督学习技术来减少人工标注的工作量。主动学习和半监督学习之间存在一定程度的互补性,因此有研究将两者联合收割机结合起来,进一步减少人工标注的工作量。然而,结合主动学习和半监督学习的SVM分类器的研究还很少见。在众多的SVM主动学习算法中,最流行的是在每次迭代中查询最接近当前分类超平面的样本的算法,本文将其表示为SVMAL。由于SVMAL只对类边界上的样本感兴趣,而忽略了对剩余大量未标记样本的利用,本文设计了一种半监督学习算法,充分利用剩余的未查询样本,进而形成一种新的主动半监督SVM算法。主动半监督SVM算法利用主动学习选择类边界样本,利用半监督学习选择类中心样本,因为类中心样本被认为能更好地描述类分布,有助于SVMAL更精确地找到边界样本。为了避免在探索类中心样本时引入过多的标签错误,使用标签变化率来保证预测标签的可靠性。实验结果表明,主动半监督SVM算法的性能明显优于纯SVM主动学习算法,从而可以进一步减少人工标注的工作量。(C)2013 Elsevier B.V.保留所有权利,
One key issue for most classification algorithms is that they need large amounts of labeled samples to train the classifier. Since manual labeling is time consuming, researchers have proposed technologies of active learning and semi-supervised learning to reduce manual labeling workload. There is a certain degree of complementarity between active learning and semi-supervised learning, and therefore some researches combine them to further reduce manual labeling workload. However, researches on combining active learning and semi-supervised learning for SVM classifier are rare. Of numerous SVM active learning algorithms, the most popular is the one that queries the sample closest to the current classification hyperplane in each iteration, which is denoted as SVMAL in this paper. Realizing that SVMAL is only interested in samples that are more likely to be on the class boundary, while ignoring the usage of the rest large amounts of unlabeled samples, this paper designs a semi-supervised learning algorithm to make full use of the rest non-queried samples, and further forms a new active semi-supervised SVM algorithm. The proposed active semi-supervised SVM algorithm uses active learning to select class boundary samples, and semi-supervised learning to select class central samples, for class central samples are believed to better describe the class distribution, and to help SVMAL finding the boundary samples more precisely. In order not to introduce too many labeling errors when exploring class central samples, the label changing rate is used to ensure the reliability of the predicted labels. Experimental results show that the proposed active semi-supervised SVM algorithm performs much better than the pure SVM active learning algorithm, and thus can further reduce manual labeling workload. (C) 2013 Elsevier B.V. All rights reserved,