SKYPEER: Efficient Subspace Skyline Computation over Distributed Data

SKYPEER: Efficient Subspace Skyline Computation over Distributed Data
复制标题

DOI:
10.1109/icde.2007.367887
复制
发表时间:
2007-04
期刊:
2007 IEEE 23rd International Conference on Data Engineering
影响因子:
--
通讯作者:
A. Vlachou;C. Doulkeridis;Y. Kotidis;M. Vazirgiannis
A. Vlachou;C. Doulkeridis;Y. Kotidis;M. Vazirgiannis
中科院分区:
其他
文献类型:
--
作者:
A. Vlachou;C. Doulkeridis;Y. Kotidis;M. Vazirgiannis

文献摘要

被引文献

相似文献

最近,Skyline查询处理受到了相当大的关注。skyline查询主要用于在多维数据集中找到一组非支配数据点。虽然大多数以前的工作都假设了一个集中的设置,在本文中,我们解决了大规模的对等(P2P)网络,其中数据集是水平分布在同行的子空间天际线查询的有效计算。依靠一个超级对等体系结构,我们提出了一个基于阈值的算法,称为SKYPEER,转发的天际线查询请求之间的同行,在这样一种方式,传输的数据量显着减少。为了有效的子空间天际线处理,我们通过定义扩展的天际线集,它包含了所有的数据元素,是必要的回答在任何任意子空间的天际线查询的支配的概念。我们证明,我们的算法提供了准确的答案,我们提出了优化技术,以减少通信成本和执行时间。最后,我们提供了一个广泛的实验评估表明,SKYPEER有效地执行,并提供了一个可行的解决方案时,需要很大程度的分布。
Skyline query processing has received considerable attention in the recent past. Mainly, the skyline query is used to find a set of non dominated data points in a multidimensional dataset. While most previous work has assumed a centralized setting, in this paper we address the efficient computation of subspace skyline queries in large-scale peer-to-peer (P2P) networks, where the dataset is horizontally distributed across the peers. Relying on a super-peer architecture we propose a threshold based algorithm, called SKYPEER, which forwards the skyline query requests among peers, in such a way that the amount of transferred data is significantly reduced. For efficient subspace skyline processing, we extend the notion of domination by defining the extended skyline set, which contains all data elements that are necessary to answer a skyline query in any arbitrary subspace. We prove that our algorithm provides the exact answers and we present optimization techniques to reduce communication cost and execution time. Finally, we provide an extensive experimental evaluation showing that SKYPEER performs efficiently and provides a viable solution when a large degree of distribution is required.