Scaling multiple-source entity resolution using statistically efficient transfer learning

Scaling multiple-source entity resolution using statistically efficient transfer learning
复制标题

DOI:
10.1145/2396761.2398606
复制
发表时间:
2012-08
期刊:
Proceedings of the 21st ACM international conference on Information and knowledge management
影响因子:
--
通讯作者:
S. Negahban;Benjamin I. P. Rubinstein;J. Gemmell
S. Negahban;Benjamin I. P. Rubinstein;J. Gemmell
中科院分区:
其他
文献类型:
--
作者:
S. Negahban;Benjamin I. P. Rubinstein;J. Gemmell

文献摘要

被引文献

相似文献

我们考虑了几乎所有将实体分辨率(ER)扩展到多个数据源的方法所面临的一个严重的,以前未探索的挑战:为每对源的相似性分数的监督学习标记训练数据的高昂成本。虽然有丰富的文献描述了成对ER的几乎所有方面,但由于从在线来源获取和存储数据的前所未有的能力,对ER驱动的功能(如丰富的搜索垂直)的兴趣,以及每个来源的噪声和缺失数据特征的独特性,现在正在出现这种新的挑战。我们在真实世界和合成数据上表明,对于最先进的技术,异构源的现实意味着标记训练数据的数量必须在源的数量中二次扩展,以保持恒定的精度/召回率。我们通过一种全新的迁移学习算法来应对这一挑战,该算法需要更少的训练数据(或者等效地,使用相同的数据实现上级准确度),并使用快速凸优化进行训练。我们的方法背后的直觉是自适应地共享结构,了解一个评分问题与所有其他评分问题共享一个共同的数据源。我们证明了我们的理论动机的方法改进了现有的多源ER技术。
We consider a serious, previously-unexplored challenge facing almost all approaches to scaling up entity resolution (ER) to multiple data sources: the prohibitive cost of labeling training data for supervised learning of similarity scores for each pair of sources. While there exists a rich literature describing almost all aspects of pairwise ER, this new challenge is arising now due to the unprecedented ability to acquire and store data from online sources, interest in features driven by ER such as enriched search verticals, and the uniqueness of noisy and missing data characteristics for each source. We show on real-world and synthetic data that for state-of-the-art techniques, the reality of heterogeneous sources means that the number of labeled training data must scale quadratically in the number of sources, just to maintain constant precision/recall. We address this challenge with a brand new transfer learning algorithm which requires far less training data (or equivalently, achieves superior accuracy with the same data) and is trained using fast convex optimization. The intuition behind our approach is to adaptively share structure learned about one scoring problem with all other scoring problems sharing a data source in common. We demonstrate that our theoretically-motivated approach improves upon existing techniques for multi-source ER.