Parallel Aggregation Queries over Star Schema: A Hierarchical Encoding Scheme and Efficient Percentile Computing as a Case
Parallel Aggregation Queries over Star Schema: A Hierarchical Encoding Scheme and Efficient Percentile Computing as a Case
复制标题
DOI:
10.1109/ispa.2011.34
复制
发表时间:
2011-05
期刊:
影响因子:
--
通讯作者:
Xiongpai Qin;Huijui Wang;Xiaoyong Du;Shan Wang
中科院分区:
文献类型:
--
作者:
Xiongpai Qin;Huijui Wang;Xiaoyong Du;Shan Wang
Big data analysis is a main challenge we meet recently. Cloud computing is attracting more and more big data analysis applications, due to its well scalability and fault-tolerance. Some aggregation functions, like SUM, can be computed in parallel, because they satisfy distributive law of addition. Unfortunately, some of statistical functions are not naturally parallelizable. That means they do not satisfy distributive law of addition. In this paper, we focus on percentile computing problem. We proposed an iterative-style prediction-based parallel algorithm in a distributed system. Prediction is done through a sampling technique. Experiment results verify the efficiency of our algorithm.