Employing unlabeled data to improve the classification performance of SVM, and its application in audio event classification

Employing unlabeled data to improve the classification performance of SVM, and its application in audio event classification
复制标题

利用无标签数据提高SVM分类性能及其在音频事件分类中的应用

DOI:
10.1016/j.knosys.2016.01.029
复制
发表时间:
2016
影响因子:
8.8
通讯作者:
Li Dengwang
Li Dengwang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Leng Yan;Sun Chengli;Xu Xinyan;Yuan Qi;Xing Shuning;Wan Honglin;Wang Jingjing;Li Dengwang

文献摘要

被引文献

相似文献

在许多分类情况下,标记样本很难获得。然而,未标记的样品很容易获得。主动学习技术可以用来解决标记问题。在众多的人工智能算法中,对支持向量机边缘带内的未标记样本进行标记是减少人工标记工作量的一种有效方法。人工智能需要人的参与,但人所能提供的时间和精力往往是有限的。因此,基于AL技术的样品标记存在很大的限制。为此,本研究的动机是对人工记忆后的加工过程进行研究。对于AL算法,其重点是探索未标记的样本的边缘带的SVM,停止后,我们的目标是调查是否可以继续探索这些未标记的样本的半监督学习(SSL)或没有。设计这样的SSL算法,一个挑战是如何计算出未标记样本的置信度,然后选择置信度高的样本。本文提出了3个确定置信度的准则:1)光滑性假设; 2)探索的正样本和探索的负样本应尽可能分别与标记的正样本和标记的负样本相似;第三条探索的阳性样品和探索的阴性样品应与标记的阴性样品和标记的阳性样品不同尽可能多的样品。基于这3个准则,本文提出了一种SSL算法--SSL_3C。在此基础上,我们将SSL_3C应用于音频事件分类领域,并在两个公开数据集上进行了实验。实验结果表明,SSL_3C能有效地提高AL处理后的分类性能。所选的未标记样本不仅置信度高,而且信息量大。此外,SSL_3C对标记和未标记训练集的大小不敏感。本文的贡献主要体现在两个方面:第一,针对SVM边缘带内的未标记样本,提出了一种有效的SSL算法来挖掘它们;第二,创新性地提出了3个判断未标记样本置信度的准则。基于这3个准则,所探索的未标记样本不仅具有较高的置信度,而且信息量很大。由于许多分类领域都存在标注问题,而SSL_3C能有效地减少人工标注的工作量,因此SSL_3C在其他领域也有着广泛的应用前景。
In many classification cases, the labeled samples are difficult to acquire. However, the unlabeled samples are easy to obtain. Active learning (AL) technology can be used to resolve the labeling problem. Among numerous kinds of AL algorithms, the one that focuses on labeling the unlabeled samples within the margin band of SVM is an effective way to decrease manual labeling workload. AL needs human involvement, but the time and energy which human can provide is often limited. Therefore, there is a big restriction for sample labeling based on the AL technology. To this end, the motivation of this work is to do studies on the processing after the AL process. For the AL algorithm which focuses on exploring the unlabeled samples within the margin band of SVM, after it stops, we aim for investigating whether such unlabeled samples can continue to be explored by semi-supervised learning (SSL) or not. To design such SSL algorithm, one of the challenges is how to figure out unlabeled samples’ confidence, and then select the ones with high confidence. In this work, we proposed 3 criterions to determine confidence, i.e. 1) the smoothness assumption; 2) the explored positive samples and the explored negative samples should be similar to the labeled positive samples and the labeled negative samples as much as possible, respectively; 3) the explored positive samples and the explored negative samples should be different from the labeled negative samples and the labeled positive samples as much as possible, respectively. Based on these 3 criterions, a SSL algorithm—SSL_3C was proposed in this work. Furthermore, we applied SSL_3C to audio event classification field, and did experiments on two public datasets. Experimental results demonstrate that SSL_3C can improve the classification performance after the AL process effectively. The selected unlabeled samples are not only of high confidence, but also very informative. Moreover, SSL_3C is not sensitive to the size of labeled and unlabeled training set. The contributions of this work lie in two aspects: first, for the unlabeled samples within the margin band of SVM, we have proposed an effective SSL algorithm to explore them; second, we innovatively proposed 3 criterions to determine unlabeled samples’ confidence. Based on these 3 criterions, the explored unlabeled samples are not only of high confidence, but also very informative. Since labeling problem exists in many classification fields, and SSL_3C can effectively decrease manual labeling workload, then the proposed SSL_3C should find widespread applications in many other fields.