Efficient Processing of Top-k Dominating Queries on Multi-Dimensional Data

Efficient Processing of Top-k Dominating Queries on Multi-Dimensional Data
复制标题

DOI:
--
复制
发表时间:
2007-09
期刊:
--
影响因子:
--
通讯作者:
Man Lung Yiu;N. Mamoulis
Man Lung Yiu;N. Mamoulis
中科院分区:
其他
文献类型:
--
作者:
Man Lung Yiu;N. Mamoulis

文献摘要

被引文献

相似文献

top-k 主导查询返回 k 个数据对象,这些数据对象主导数据集中最多数量的对象。此查询是决策支持的重要工具,因为它为数据分析师提供了查找重要对象的直观方法。此外,它结合了top-k和skyline查询的优点,同时又避免了它们的缺点:(i)输出大小可以控制,(ii)不需要用户指定排序函数,(iii)结果与不同维度的尺度无关。尽管很重要,但 top-k 主导查询尚未得到研究界的足够关注。在本文中,我们设计了适用于索引多维数据的专用算法,并充分利用了问题的特征。对合成数据集的实验表明,我们的算法显着优于之前基于天际线的方法,而我们对真实数据集的结果显示了 top-k 主导查询的意义。
The top-k dominating query returns k data objects which dominate the highest number of objects in a dataset. This query is an important tool for decision support since it provides data analysts an intuitive way for finding significant objects. In addition, it combines the advantages of top-k and skyline queries without sharing their disadvantages: (i) the output size can be controlled, (ii) no ranking functions need to be specified by users, and (iii) the result is independent of the scales at different dimensions. Despite their importance, top-k dominating queries have not received adequate attention from the research community. In this paper, we design specialized algorithms that apply on indexed multi-dimensional data and fully exploit the characteristics of the problem. Experiments on synthetic datasets demonstrate that our algorithms significantly outperform a previous skyline-based approach, while our results on real datasets show the meaningfulness of top-k dominating queries.