Efficient Indexing Methods for Probabilistic Threshold Queries over Uncertain Data

Efficient Indexing Methods for Probabilistic Threshold Queries over Uncertain Data
复制标题

DOI:
10.1016/b978-012088469-8.50077-2
复制
发表时间:
2004-08
期刊:
--
影响因子:
--
通讯作者:
Reynold Cheng;Yuni Xia;Sunil Prabhakar;Rahul Shah;J. Vitter
Reynold Cheng;Yuni Xia;Sunil Prabhakar;Rahul Shah;J. Vitter
中科院分区:
其他
文献类型:
--
作者:
Reynold Cheng;Yuni Xia;Sunil Prabhakar;Rahul Shah;J. Vitter

文献摘要

被引文献

相似文献

传感器数据库不可能包含每个传感器在所有时间点的精确值。由于测量和取样误差以及资源限制,这种不确定性是这些系统固有的。为了避免基于陈旧数据得出错误的结论,最近有人提出使用不确定性区间,将每个数据项建模为范围和相关的概率密度函数(pdf),而不是单个值。查询这些不确定的数据将不精确性引入到答案中,以概率值的形式指定答案满足查询的可能性。这些查询比传统查询的评估成本更高,但由于答案伴随的概率,可以保证是正确的,信息量更大。虽然答案概率是有用的,但对于许多应用程序来说,只需要知道概率是否超过给定的阈值,我们称之为概率阈值(PTQ)。在本文中,我们解决这些类型的查询的有效计算。
It is infeasible for a sensor database to contain the exact value of each sensor at all points in time. This uncertainty is inherent in these systems due to measurement and sampling errors, and resource limitations. In order to avoid drawing erroneous conclusions based upon stale data, the use of uncertainty intervals that model each data item as a range and associated probability density function (pdf) rather than a single value has recently been proposed. Querying these uncertain data introduces imprecision into answers, in the form of probability values that specify the likeliness the answer satisfies the query. These queries are more expensive to evaluate than their traditional counterparts but are guaranteed to be correct and more informative due to the probabilities accompanying the answers. Although the answer probabilities are useful, for many applications, it is only necessary to know whether the probability exceeds a given threshold–we term these Probabilistic Threshold Queries (PTQ). In this paper we address the efficient computation of these types of queries.