Active transfer learning of matching query results across multiple sources

Active transfer learning of matching query results across multiple sources
复制标题

跨多个源匹配查询结果的主动迁移学习

DOI:
10.1007/s11704-015-4068-3
复制
发表时间:
2015-08
影响因子:
4.2
通讯作者:
何天旭
何天旭
中科院分区:
计算机科学3区
文献类型:
--
作者:
辛洁;崔志明;赵朋朋;何天旭

文献摘要

参考文献

被引文献

相似文献

实体解析(Entity Resolution,ER)是对同一个真实的世界对象的不同表现形式进行识别和分组的问题。已经开发出了一些数学方法,其中大多数任务在监督学习下提供了上级性能。然而,标记训练数据的高昂成本仍然是从在线源中检测重复查询记录的巨大障碍。此外,噪声数据与缺失元素的独特组合使ER任务更具挑战性。为了解决这个问题,迁移学习已经被采用来自适应地共享多个源之间的相似性评分问题的学习的共同结构。虽然这样的技术降低了标记成本,使它是线性的,相对于源的数量,其随机抽样策略是不够成功,以处理普通的样本不平衡问题。在本文中,我们提出了一种新的多源主动迁移学习框架,从所有源中联合选择较少的数据实例来训练具有恒定精度/召回率的分类器。我们的方法背后的直觉是积极标记信息量最大的样本,同时在源之间自适应地转移集体知识。通过这种方式,即使对于不平衡或质量不同的源,学习的分类器也可以是标签经济的和灵活的。我们将我们的方法与真实数据集上的最先进方法进行比较。我们的实验结果表明,我们的主动迁移学习算法可以实现令人印象深刻的性能与标记的样本少得多的记录匹配与众多和不同的来源。
Entity resolution (ER) is the problem of identifying and grouping different manifestations of the same real world object. Algorithmic approaches have been developed where most tasks offer superior performance under supervised learning. However, the prohibitive cost of labeling training data is still a huge obstacle for detecting duplicate query records from online sources. Furthermore, the unique combinations of noisy data with missing elements make ER tasks more challenging. To address this, transfer learning has been adopted to adaptively share learned common structures of similarity scoring problems between multiple sources. Although such techniques reduce the labeling cost so that it is linear with respect to the number of sources, its random sampling strategy is not successful enough to handle the ordinary sample imbalance problem. In this paper, we present a novel multi-source active transfer learning framework to jointly select fewer data instances from all sources to train classifiers with constant precision/recall. The intuition behind our approach is to actively label the most informative samples while adaptively transferring collective knowledge between sources. In this way, the classifiers that are learned can be both label-economical and flexible even for imbalanced or quality diverse sources. We compare our method with the state-of-the-art approaches on real-word datasets. Our experimental results demonstrate that our active transfer learning algorithm can achieve impressive performance with far fewer labeled samples for record matching with numerous and varied sources.
DOI: 10.1145/2396761.2398606
发表时间: 2012-08
期刊: Proceedings of the 21st ACM international conference on Information and knowledge management
影响因子: --
作者:
S. Negahban;Benjamin I. P. Rubinstein;J. Gemmell
通讯作者: S. Negahban;Benjamin I. P. Rubinstein;J. Gemmell
DOI: 10.1145/1458082.1458090
发表时间: 2008-10
期刊: --
影响因子: --
作者:
Shui-Lung Chuang;K. Chang
通讯作者: Shui-Lung Chuang;K. Chang
DOI: 10.1145/1807167.1807252
发表时间: 2010-06
期刊: Proceedings of the 2010 ACM SIGMOD International Conference on Management of data
影响因子: --
作者:
A. Arasu;M. Götz;R. Kaushik
通讯作者: A. Arasu;M. Götz;R. Kaushik
DOI: --
发表时间: 2010
期刊: Journal of Frontiers of Computer Science and Technology
影响因子: --
作者:
Liubao Wei
通讯作者: Liubao Wei
DOI: 10.1145/1645953.1646138
发表时间: 2009-11
期刊: Proceedings of the 18th ACM conference on Information and knowledge management
影响因子: --
作者:
Ming-Hay Luk;Man Lung Yiu;Eric Lo
通讯作者: Ming-Hay Luk;Man Lung Yiu;Eric Lo