Parallel Aggregation Queries over Star Schema: A Hierarchical Encoding Scheme and Efficient Percentile Computing as a Case

Parallel Aggregation Queries over Star Schema: A Hierarchical Encoding Scheme and Efficient Percentile Computing as a Case
复制标题

DOI:
10.1109/ispa.2011.34
复制
发表时间:
2011-05
期刊:
2011 IEEE Ninth International Symposium on Parallel and Distributed Processing with Applications
影响因子:
--
通讯作者:
Xiongpai Qin;Huijui Wang;Xiaoyong Du;Shan Wang
Xiongpai Qin;Huijui Wang;Xiaoyong Du;Shan Wang
中科院分区:
其他
文献类型:
--
作者:
Xiongpai Qin;Huijui Wang;Xiaoyong Du;Shan Wang

文献摘要

被引文献

相似文献

大数据分析是我们最近遇到的一个主要挑战。云计算以其良好的可扩展性和容错性吸引了越来越多的大数据分析应用。一些聚集函数,如SUM,可以并行计算,因为它们满足分配加法定律。不幸的是,一些统计函数不是自然可并行化的。这意味着它们不满足加法的分配定律。在本文中,我们主要研究百分位数计算问题。在分布式系统中,提出了一种基于迭代预测的并行算法。预测是通过抽样技术进行的。实验结果验证了该算法的有效性。
Big data analysis is a main challenge we meet recently. Cloud computing is attracting more and more big data analysis applications, due to its well scalability and fault-tolerance. Some aggregation functions, like SUM, can be computed in parallel, because they satisfy distributive law of addition. Unfortunately, some of statistical functions are not naturally parallelizable. That means they do not satisfy distributive law of addition. In this paper, we focus on percentile computing problem. We proposed an iterative-style prediction-based parallel algorithm in a distributed system. Prediction is done through a sampling technique. Experiment results verify the efficiency of our algorithm.