On-Demand Query Result Cleaning
On-Demand Query Result Cleaning
复制标题
按需查询结果清理
DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Jan Chomicki
中科院分区:
文献类型:
--
作者:
Ying Yang;Oliver Kennedy;Jan Chomicki
Incomplete data is ubiquitous. When a user issues a query over incomplete data, the results may contain incomplete data as well. If a user requires high precision query re-sults, or current estimation algorithms fail to make accurate estimates on incomplete data, data collection by humans may instead be used to find values for, or to confirm this incomplete data. We propose an approach that incrementally confirms incomplete data: First, queries on incomplete data are processed by a probabilistic database system. In-complete data in the query results is represented in a form called candidate questions . Second, we incrementally solicit user feedback to confirm candidate questions. The challenge of this approach is to determine in what order to confirm candidate questions with the user. To solve this, we design a framework for ranking candidate questions for user confirmation using a concept that we call cost of perfect information (CPI). The core component of CPI is a penalty function that is based on entropy . We compare each candidate question’s CPI and choose an optimal candidate question to solicit user feedback. Our approach achieves accurate query results with low confirmation and computation costs. Experiments on a real dataset show that our approach outperforms other strategies.