On-Demand Query Result Cleaning

On-Demand Query Result Cleaning
复制标题

按需查询结果清理

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Jan Chomicki
Jan Chomicki
中科院分区:
--
文献类型:
--
作者:
Ying Yang;Oliver Kennedy;Jan Chomicki

文献摘要

被引文献

相似文献

不完整的数据无处不在。当用户对不完整数据发出查询时,结果也可能包含不完整数据。如果用户需要高精度的查询结果,或者现有的估计算法不能对不完整数据做出准确的估计,则可以使用人工收集的数据来对这些不完整数据进行find值或firm。我们提出了一种增量控制fiRMS不完全数据的方法:首先,由概率数据库系统处理对不完整数据的查询。查询结果中的不完整数据以称为候选问题的形式表示。其次,我们逐步征求用户反馈,以应对fiRM候选问题。这种方法的挑战是确定以什么顺序向用户提出fiRM候选问题。为了解决这一问题,我们设计了一个框架,用完美信息成本(CPI)的概念对候选问题进行排序。CPI的核心成分是基于熵的惩罚函数。我们比较每个候选问题的CPI,并选择一个最优的候选问题来征求用户反馈。该方法以较低的构造和计算代价获得了准确的查询结果。在真实数据集上的实验表明,该方法的性能优于其他策略。
Incomplete data is ubiquitous. When a user issues a query over incomplete data, the results may contain incomplete data as well. If a user requires high precision query re-sults, or current estimation algorithms fail to make accurate estimates on incomplete data, data collection by humans may instead be used to find values for, or to confirm this incomplete data. We propose an approach that incrementally confirms incomplete data: First, queries on incomplete data are processed by a probabilistic database system. In-complete data in the query results is represented in a form called candidate questions . Second, we incrementally solicit user feedback to confirm candidate questions. The challenge of this approach is to determine in what order to confirm candidate questions with the user. To solve this, we design a framework for ranking candidate questions for user confirmation using a concept that we call cost of perfect information (CPI). The core component of CPI is a penalty function that is based on entropy . We compare each candidate question’s CPI and choose an optimal candidate question to solicit user feedback. Our approach achieves accurate query results with low confirmation and computation costs. Experiments on a real dataset show that our approach outperforms other strategies.