Incorporating Diversity in Active Learning with Support Vector Machines

Incorporating Diversity in Active Learning with Support Vector Machines
复制标题

DOI:
--
复制
发表时间:
2003-08
期刊:
--
影响因子:
--
通讯作者:
K. Brinker
K. Brinker
中科院分区:
其他
文献类型:
--
作者:
K. Brinker

文献摘要

被引文献

相似文献

在许多真实的应用中,主动选择训练样本可以显著减少用于学习分类函数的标记训练样本的数量。在支持向量机领域已经提出了不同的策略,迭代地从一组未标记的例子中选择一个新的例子,查询相应的类标签,然后执行当前分类器的再训练。然而,为了减少训练的计算时间,可能需要选择批量的新训练示例而不是单个示例。单个示例的策略可以直接扩展到选择批次,方法是选择h > 1的示例,这些示例对于单个选择标准具有最高值。我们提出了一种新的方法,是专门设计来构建批次,并采用了多样性措施。它具有较低的计算要求,使其成为可行的大规模问题与数千个例子。实验结果表明,这种方法提供了一种更快的方法,以达到一定程度的泛化精度的标记的例子的数量。
In many real world applications, active selection of training examples can significantly reduce the number of labelled training examples to learn a classification function. Different strategies in the field of support vector machines have been proposed that iteratively select a single new example from a set of unlabelled examples, query the corresponding class label and then perform retraining of the current classifier. However, to reduce computational time for training, it might be necessary to select batches of new training examples instead of single examples. Strategies for single examples can be extended straightforwardly to select batches by choosing the h > 1 examples that get the highest values for the individual selection criterion. We present a new approach that is especially designed to construct batches and incorporates a diversity measure. It has low computational requirements making it feasible for large scale problems with several thousands of examples. Experimental results indicate that this approach provides a faster method to attain a level of generalization accuracy in terms of the number of labelled examples.