Improving SVM accuracy by training on auxiliary data sources

Improving SVM accuracy by training on auxiliary data sources
复制标题

DOI:
10.1145/1015330.1015436
复制
发表时间:
2004-07
期刊:
Proceedings of the twenty-first international conference on Machine learning
影响因子:
--
通讯作者:
Pengcheng Wu;Thomas G. Dietterich
Pengcheng Wu;Thomas G. Dietterich
中科院分区:
其他
文献类型:
--
作者:
Pengcheng Wu;Thomas G. Dietterich

文献摘要

被引文献

相似文献

监督学习的标准模型假设训练和测试数据来自相同的底层分布。本文探讨了一个应用程序,其中的第二个,辅助,数据源是从不同的分布绘制。这种辅助数据比训练和测试数据更丰富,但质量明显较低。在SVM框架中,训练样本有两个角色:(a)作为约束学习过程的数据点,以及(B)作为可以形成分类器定义的一部分的候选支持向量。本文考虑在这两种作用中使用辅助数据。这个辅助数据框架被应用到一个问题,枫树和橡树的叶子的图像分类使用的内核来自叶子的形状。实验表明,当训练数据集非常小时,使用辅助数据进行训练可以大大提高准确率,即使辅助数据与训练(和测试)数据显著不同。本文还介绍了调整辅助数据点的核分数的技术,使它们更具有可比性的训练数据点。
The standard model of supervised learning assumes that training and test data are drawn from the same underlying distribution. This paper explores an application in which a second, auxiliary, source of data is available drawn from a different distribution. This auxiliary data is more plentiful, but of significantly lower quality, than the training and test data. In the SVM framework, a training example has two roles: (a) as a data point to constrain the learning process and (b) as a candidate support vector that can form part of the definition of the classifier. The paper considers using the auxiliary data in either (or both) of these roles. This auxiliary data framework is applied to a problem of classifying images of leaves of maple and oak trees using a kernel derived from the shapes of the leaves. Experiments show that when the training data set is very small, training with auxiliary data can produce large improvements in accuracy, even when the auxiliary data is significantly different from the training (and test) data. The paper also introduces techniques for adjusting the kernel scores of the auxiliary data points to make them more comparable to the training data points.