Reporting l most influential objects in uncertain databases based on probabilistic reverse top-k queries
Reporting l most influential objects in uncertain databases based on probabilistic reverse top-k queries
复制标题
基于概率反向top-k查询报告不确定数据库中最有影响力的对象
DOI:
10.1016/j.ins.2017.04.028
复制
发表时间:
2017-09-01
影响因子:
8.1
通讯作者:
Li, Keqin
中科院分区:
文献类型:
--
作者:
Xiao, Guoqing;Li, Kenli;Li, Keqin
Reverse top-k queries are proposed from the perspective of a product manufacturer, which are essential for manufacturers to assess the potential market. However, the existing approaches for reverse top-k queries are all based on the assumption that the underlying data are exact (or certain). Due to the intrinsic differences between uncertain and certain data, these methods cannot be applied to process uncertain data sets directly. Motivated by this, in this paper, we firstly model the probabilistic reverse top-k queries over uncertain data. Moreover, we formulate a probabilistic top-l influential query, that reports the 1 most influential objects having the largest impact factors, where the impact factor of an object is defined as the cardinality of its probabilistic reverse top-k query result set. We present effective pruning heuristics for speeding up the queries. Particularly, we exploit several properties of probabilistic threshold top-k queries and probabilistic skyline queries to reduce the search space of this problem. In addition, an upper bound of the potential users is estimated to reduce the cost of computing the probabilistic reverse top-k queries for the candidate objects. Finally, efficient query algorithms are presented seamlessly with integration of the proposed pruning strategies. Extensive experiments using both real-world and synthetic data sets demonstrate the efficiency and effectiveness of our proposed algorithms. (C) 2017 Elsevier Inc. All rights reserved.