Cleaning uncertain data with quality guarantees

Cleaning uncertain data with quality guarantees
复制标题

DOI:
10.14778/1453856.1453935
复制
发表时间:
2008-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Reynold Cheng;Jinchuan Chen;Xike Xie
Reynold Cheng;Jinchuan Chen;Xike Xie
中科院分区:
其他
文献类型:
--
作者:
Reynold Cheng;Jinchuan Chen;Xike Xie

文献摘要

被引文献

相似文献

不确定或不精确的数据在基于位置的服务、传感器监控以及数据收集和集成等应用中普遍存在。对于这些应用,可以使用概率数据库来存储不确定的数据,并提供查询设施来产生具有统计置信度的答案。假设有有限数量的资源可用于对数据库进行“清理”(例如,通过探测一些传感器数据值以获得它们的最新值),我们解决了选择要清理的不确定对象集的问题,以便在查询答案的质量方面实现最佳改进。为此,我们提出了PWS-Quality度量,这是一种在可能世界语义下量化查询答案歧义性的通用度量。我们研究了如何针对两个主要查询类别有效地评估PWS质量:(1)独立于其他元组检查元组可满足性的查询(例如,范围查询);以及(2)需要知道元组相对排名的查询(例如,MAX查询)。然后,我们提出了一个多项式时间解,以实现PWS质量的最优改善。还给出了其他快速启发式算法。在真实数据集和合成数据集上进行的实验表明,PWS质量度量可以快速评估,并且我们的清理算法提供了高效的最优解。据我们所知,这是第一项为概率数据库开发质量度量并调查如何将此类度量用于数据清理目的的工作。
Uncertain or imprecise data are pervasive in applications like location-based services, sensor monitoring, and data collection and integration. For these applications, probabilistic databases can be used to store uncertain data, and querying facilities are provided to yield answers with statistical confidence. Given that a limited amount of resources is available to "clean" the database (e.g., by probing some sensor data values to get their latest values), we address the problem of choosing the set of uncertain objects to be cleaned, in order to achieve the best improvement in the quality of query answers. For this purpose, we present the PWS-quality metric, which is a universal measure that quantifies the ambiguity of query answers under the possible world semantics. We study how PWS-quality can be efficiently evaluated for two major query classes: (1) queries that examine the satisfiability of tuples independent of other tuples (e.g., range queries); and (2) queries that require the knowledge of the relative ranking of the tuples (e.g., MAX queries). We then propose a polynomial-time solution to achieve an optimal improvement in PWS-quality. Other fast heuristics are presented as well. Experiments, performed on both real and synthetic datasets, show that the PWS-quality metric can be evaluated quickly, and that our cleaning algorithm provides an optimal solution with high efficiency. To our best knowledge, this is the first work that develops a quality metric for a probabilistic database, and investigates how such a metric can be used for data cleaning purposes.