Improving the Quality of Web-Based Data Imputation With Crowd Intervention

Improving the Quality of Web-Based Data Imputation With Crowd Intervention
复制标题

通过群体干预提高基于网络的数据插补的质量

DOI:
10.1109/tkde.2019.2954087
复制
发表时间:
2021-06
期刊:
IEEE TKDE (CCF-A类期刊)
影响因子:
--
通讯作者:
Xiaofang Zhou
Xiaofang Zhou
中科院分区:
其他
文献类型:
--
作者:
Binbin Gu;Zhixu Li;An Liu;Jiajie Xu;Lei Zhao;Xiaofang Zhou

文献摘要

参考文献

相似文献

数据不完整性是数据库中常见的数据质量问题。最近的工作提出了检索丢失的字符串值从万维网更高的插补召回,但另一方面,需要引入网络噪声的插补结果的风险。由于质量模型的复杂性和缺乏足够的真实数据,到目前为止,还没有一种有效的方法来控制基于Web的数据填补的质量。本文首先建立了一个基于EM的网络数据插补质量模型,该模型综合考察了三个关键因素,Web源的精确度、Web源之间的相关性以及所采用的提取器的精确度和召回率。然而,基于EM的质量模型的准确性可能会受到损害时,EM(期望最大化)的假设,“大多数人同意的真理”并不成立,在某些情况下。为了解决这个问题,我们引入了人群干预来帮助改进质量模型。虽然一个简单但昂贵的方法是让人群识别所有这些不期望的情况,并为这些空白提供正确的插补值,但最具人群经济性的方法是选择一小组空白进行基于人群的插补,其结果有助于调整基于EM的质量模型。为了实现这一点,提出了一种自适应的空白选择策略,以选择一系列的空白人群为基础的填补。此外,我们致力于寻找适当的时间停止进一步的人群干预,以平衡人群效率和质量改善。我们在三个真实的世界和一个模拟数据集上进行的实验证明,所提出的质量模型可以有效地帮助提高基于Web的插补结果的质量超过15%,而我们的人群成本节省策略节省超过75%的人群成本。
Data incompleteness is a common data quality problem in databases. Recent work proposes to retrieve missing string values from the World Wide Web for higher imputation recall, but on the other hand, takes the risk of introducing web noises into the imputation results. So far there lacks an effective way to control the quality of web-based data imputation, given the complexity of the quality model and lacking of enough ground truth data. In this article, an EM-based quality model is first built for web-based data imputation which investigates three key factors jointly, i.e., precision of web sources, correlation among web sources, and precision and recall of the employed extractors. However, the accuracy of the EM-based quality model could be harmed when the EM (Expectation Maximization) assumption that “the majority agree on the truth” does not hold in some cases. To solve this problem, we introduce crowd intervention to help improve the quality model. While a straightforward but expensive way is to let the crowd to identify all these undesirable cases and provide the right imputation values for these blanks, a most crowd-economic way is to select a small set of blanks for crowd-based imputation, whose results could help to adjust the EM-based quality model towards a better one. To achieve this, an adaptive blank selection strategy is proposed to select a sequence of blanks for crowd-based imputation. Also, we work on finding a proper time to stop further crowd intervention for the balance of crowd efficiency and quality improvement. Our experiments performed on three real world and one simulated data collections prove that the proposed quality model can effectively help improve the quality of the web-based imputation results by more than 15 percent, while our crowd cost saving strategy saves more than 75 percent crowd cost.
DOI: 10.1214/aos/1028674845
发表时间: 2002-06
影响因子: 4.5
作者:
Qihua Wang;J. Rao
通讯作者: Qihua Wang;J. Rao
DOI: 10.1145/1989323.1989331
发表时间: 2011-06
期刊: IEEE/RSJ International Conference on Intelligent Robots and Systems
影响因子: --
作者:
M. Franklin;Donald Kossmann;Tim Kraska;Sukriti Ramesh;Reynold Xin
通讯作者: M. Franklin;Donald Kossmann;Tim Kraska;Sukriti Ramesh;Reynold Xin
DOI: 10.1145/2939672.2939816
发表时间: 2016-08
期刊: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子: --
作者:
Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu
通讯作者: Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu
DOI: 10.1186/s12859-014-0346-6
发表时间: 2014-11-05
期刊: BMC bioinformatics
影响因子: 3
作者:
Liao SG;Lin Y;Kang DD;Chandra D;Bon J;Kaminski N;Sciurba FC;Tseng GC
通讯作者: Tseng GC
DOI: --
发表时间: 2006
期刊: --
影响因子: --
作者:
A. Rogier;T. Donders;T. Stijnen
通讯作者: A. Rogier;T. Donders;T. Stijnen