Improving the Quality of Web-Based Data Imputation With Crowd Intervention
Improving the Quality of Web-Based Data Imputation With Crowd Intervention
复制标题
通过群体干预提高基于网络的数据插补的质量
DOI:
10.1109/tkde.2019.2954087
复制
发表时间:
2021-06
期刊:
影响因子:
--
通讯作者:
Xiaofang Zhou
中科院分区:
文献类型:
--
作者:
Binbin Gu;Zhixu Li;An Liu;Jiajie Xu;Lei Zhao;Xiaofang Zhou
Data incompleteness is a common data quality problem in databases. Recent work proposes to retrieve missing string values from the World Wide Web for higher imputation recall, but on the other hand, takes the risk of introducing web noises into the imputation results. So far there lacks an effective way to control the quality of web-based data imputation, given the complexity of the quality model and lacking of enough ground truth data. In this article, an EM-based quality model is first built for web-based data imputation which investigates three key factors jointly, i.e., precision of web sources, correlation among web sources, and precision and recall of the employed extractors. However, the accuracy of the EM-based quality model could be harmed when the EM (Expectation Maximization) assumption that “the majority agree on the truth” does not hold in some cases. To solve this problem, we introduce crowd intervention to help improve the quality model. While a straightforward but expensive way is to let the crowd to identify all these undesirable cases and provide the right imputation values for these blanks, a most crowd-economic way is to select a small set of blanks for crowd-based imputation, whose results could help to adjust the EM-based quality model towards a better one. To achieve this, an adaptive blank selection strategy is proposed to select a sequence of blanks for crowd-based imputation. Also, we work on finding a proper time to stop further crowd intervention for the balance of crowd efficiency and quality improvement. Our experiments performed on three real world and one simulated data collections prove that the proposed quality model can effectively help improve the quality of the web-based imputation results by more than 15 percent, while our crowd cost saving strategy saves more than 75 percent crowd cost.
登录
查看更多内容
影响因子:
4.5
作者:
Qihua Wang;J. Rao
通讯作者:
Qihua Wang;J. Rao
DOI:
10.1145/1989323.1989331
发表时间:
2011-06
期刊:
IEEE/RSJ International Conference on Intelligent Robots and Systems
影响因子:
--
作者:
M. Franklin;Donald Kossmann;Tim Kraska;Sukriti Ramesh;Reynold Xin
通讯作者:
M. Franklin;Donald Kossmann;Tim Kraska;Sukriti Ramesh;Reynold Xin
DOI:
10.1145/2939672.2939816
发表时间:
2016-08
期刊:
Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子:
--
作者:
Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu
通讯作者:
Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu
影响因子:
3
作者:
Liao SG;Lin Y;Kang DD;Chandra D;Bon J;Kaminski N;Sciurba FC;Tseng GC
通讯作者:
Tseng GC
DOI:
--
发表时间:
2006
期刊:
--
影响因子:
--
作者:
A. Rogier;T. Donders;T. Stijnen
通讯作者:
A. Rogier;T. Donders;T. Stijnen