Top-k Dominating Queries on Incomplete Data

Top-k Dominating Queries on Incomplete Data
复制标题

对不完整数据的 Top-k 主导查询

DOI:
10.1109/tkde.2015.2460742
复制
发表时间:
2016-01-01
影响因子:
8.9
通讯作者:
Cui, Huiyong
Cui, Huiyong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Miao, Xiaoye;Gao, Yunjun;Cui, Huiyong

文献摘要

被引文献

相似文献

top-k dominating(TKD)查询返回在给定数据集中占主导地位的最大数量的对象的k个对象。它结合了skyline和top-k查询的优点,在许多决策支持应用中发挥着重要作用。由于设备故障、隐私保护、数据丢失等原因,不完整数据广泛存在于真实的数据集中,本文首次对不完整数据(包括维数值缺失的数据)的TKD查询进行了系统的研究.我们正式这个问题,并提出了一套有效的算法,回答TKD查询不完整的数据。我们的方法采用了一些新的技术,如上限分数修剪,位图修剪,部分分数修剪,以提高查询效率。使用真实的和合成数据集的广泛的实验评估表明,我们开发的修剪算法的有效性和我们提出的算法的性能。
The top-k dominating (TKD) query returns the k objects that dominate the maximum number of objects in a given dataset. It combines the advantages of skyline and top-k queries, and plays an important role in many decision support applications. Incomplete data exists in a wide spectrum of real datasets, due to device failure, privacy preservation, data loss, and so on. In this paper, for the first time, we carry out a systematic study of TKD queries on incomplete data, which involves the data having some missing dimensional value(s). We formalize this problem, and propose a suite of efficient algorithms for answering TKD queries over incomplete data. Our methods employ some noveltechniques, such as upper bound score pruning, bitmap pruning, and partial score pruning, to boost query efficiency. Extensive experimental evaluation using both real and synthetic datasets demonstrates the effectiveness of our developed pruning heuristics and the performance of our presented algorithms.