BOOSTED UNSUPERVISED MULTI-SOURCE SELECTION FOR DOMAIN ADAPTATION

BOOSTED UNSUPERVISED MULTI-SOURCE SELECTION FOR DOMAIN ADAPTATION
复制标题

DOI:
10.5194/isprs-annals-iv-1-w1-229-2017
复制
发表时间:
2017-05
期刊:
ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences
影响因子:
--
通讯作者:
Karsten Vogt;A. Paul;J. Ostermann;F. Rottensteiner;C. Heipke
Karsten Vogt;A. Paul;J. Ostermann;F. Rottensteiner;C. Heipke
中科院分区:
其他
文献类型:
--
作者:
Karsten Vogt;A. Paul;J. Ostermann;F. Rottensteiner;C. Heipke

文献摘要

相似文献

抽象。监督机器学习需要高质量、密集采样和标记的训练数据。迁移学习(TL)技术已经被设计为通过使在不同但相关的(源)训练数据上训练的分类器适应新的(目标)数据集来减少这种依赖性。TL中的一个问题是如何快速和鲁棒地量化源的相关性,因为从不相关的数据中转移知识会降低分类器的性能。在本文中,我们提出了一种方法,可以选择一个接近最佳的源从大量的候选源。这种操作仅依赖于数据的边缘概率分布,从而允许使用通常丰富的未标记数据。我们将此方法扩展到多源选择,通过优化加权组合的来源。使用非常快速的提升式优化方案计算源权重。我们的方法的运行时的复杂度线性缩放的候选源的数量和训练集的大小,因此适用于非常大的数据集。我们还提出了一个修改现有的TL算法来处理多个加权训练集。我们的方法进行了评估五个调查区域。实验表明,我们的源选择方法在区分相关和不相关的源是有效的,几乎总是产生的结果在3%的整体精度的分类器的基础上完全标记的训练数据。我们还表明,使用选定的源作为TL方法的训练数据将额外导致性能提高。
Abstract. Supervised machine learning needs high quality, densely sampled and labelled training data. Transfer learning (TL) techniques have been devised to reduce this dependency by adapting classifiers trained on different, but related, (source) training data to new (target) data sets. A problem in TL is how to quantify the relatedness of a source quickly and robustly, because transferring knowledge from unrelated data can degrade the performance of a classifier. In this paper, we propose a method that can select a nearly optimal source from a large number of candidate sources. This operation depends only on the marginal probability distributions of the data, thus allowing the use of the often abundant unlabelled data. We extend this method to multi-source selection by optimizing a weighted combination of sources. The source weights are computed using a very fast boosting-like optimization scheme. The run-time complexity of our method scales linearly in regard to the number of candidate sources and the size of the training set and is thus applicable to very large data sets. We also propose a modification of an existing TL algorithm to handle multiple weighted training sets. Our method is evaluated on five survey regions. The experiments show that our source selection method is effective in discriminating between related and unrelated sources, almost always generating results within 3% in overall accuracy of a classifier based on fully labelled training data. We also show that using the selected source as training data for a TL method will additionally result in a performance improvement.