Support vector machine active learning with applications to text classification

Support vector machine active learning with applications to text classification
复制标题

DOI:
10.1162/153244302760185243
复制
发表时间:
2002-12-01
影响因子:
6
通讯作者:
Koller, D
Koller, D
中科院分区:
计算机科学3区
文献类型:
--
作者:
Tong, S;Koller, D

文献摘要

被引文献

相似文献

支持向量机在许多现实世界的学习任务中取得了显著的成功。然而,与大多数机器学习算法一样,它们通常使用事先分类的随机选择的训练集来应用。在很多情况下,我们也可以选择使用基于池的主动学习。而不是使用随机选择的训练集,学习者可以访问一个未标记的实例池,并可以请求其中一些实例的标签。我们引入了一种用支持向量机进行主动学习的新算法,即选择下一步请求哪个实例的算法。我们使用版本空间的概念为算法提供了理论动机。我们提出的实验结果表明,采用我们的主动学习方法可以显着减少在标准归纳和转换设置中对标记训练实例的需求。
Support vector machines have met with significant success in numerous real-world learning tasks. However, like most machine learning algorithms, they are generally applied using a randomly selected training set classified in advance. In many settings, we also have the option of using pool-based active learning. Instead of using a randomly selected training set, the learner has access to a pool of unlabeled instances and can request the labels for some number of them. We introduce a new algorithm for performing active learning with support vector machines, i.e., an algorithm for choosing which instances to request next. We provide a theoretical motivation for the algorithm using the notion of a version space. We present experimental results showing that employing our active learning method can significantly reduce the need for labeled training instances in both the standard inductive and transductive settings.