Human-in-the-loop Data Integration

Human-in-the-loop Data Integration
复制标题

DOI:
10.14778/3137765.3137833
复制
发表时间:
2017-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Guoliang Li
Guoliang Li
中科院分区:
其他
文献类型:
--
作者:
Guoliang Li

文献摘要

被引文献

相似文献

数据集成的目的是将不同来源的数据进行集成,为用户提供统一的视图。然而,单纯的自动化方法无法完全解决数据集成问题。我们提出了一个混合的人机数据集成框架,利用人类的能力来解决这个问题,并将其初步应用于实体匹配问题。该框架首先使用基于规则的算法来识别可能的匹配对,然后利用人群对这些候选对进行细化,以计算实际的匹配对。在第一步中,我们提出基于相似度的规则和基于知识的规则来获得一些候选匹配对,并基于给定的正反例开发有效的算法来学习这些规则。我们构建了一个分布式内存系统DIMA来有效地应用这些规则。在第二步中,我们提出了一个选择-推理-改进框架,该框架使用人群来验证候选对。我们首先选择一些“有益”的任务向人群提问,然后根据被提问任务的众包结果,利用及物性和偏序来推断未被提问任务的答案。接下来,我们提炼了由于人群的不一致而导致的高不确定性的推断答案。我们开发了一个众包数据库系统CDB,并将其部署在真正的众包平台上。CDB允许用户使用类似sql的语言来处理基于人群的查询。最后,我们提出了人在环数据集成方面的新挑战。
Data integration aims to integrate data in different sources and provide users with a unified view. However, data integration cannot be completely addressed by purely automated methods. We propose a hybrid human-machine data integration framework that harnesses human ability to address this problem, and apply it initially to the problem of entity matching. The framework first uses rule-based algorithms to identify possible matching pairs and then utilizes the crowd to refine these candidate pairs in order to compute actual matching pairs. In the first step, we propose similarity-based rules and knowledge-based rules to obtain some candidate matching pairs, and develop effective algorithms to learn these rules based on some given positive and negative examples. We build a distributed in-memory system DIMA to efficiently apply these rules. In the second step, we propose a selection-inference-refine framework that uses the crowd to verify the candidate pairs. We first select some "beneficial" tasks to ask the crowd and then use transitivity and partial order to infer the answers of unasked tasks based on the crowdsourcing results of the asked tasks. Next we refine the inferred answers with high uncertainty due to the disagreement from the crowd. We develop a crowd-powered database system CDB and deploy it on real crowdsourcing platforms. CDB allows users to utilize a SQL-like language for processing crowd-based queries. Lastly, we provide emerging challenges in human-in-the-loop data integration.