P-DOT: A model of computation for big data

P-DOT: A model of computation for big data
复制标题

DOI:
10.1080/17445760.2015.1016515
复制
发表时间:
2013-12
期刊:
2013 IEEE International Conference on Big Data
影响因子:
--
通讯作者:
Tao Luo;Yin Liao;Guoliang Chen;Yunquan Zhang
Tao Luo;Yin Liao;Guoliang Chen;Yunquan Zhang
中科院分区:
其他
文献类型:
--
作者:
Tao Luo;Yin Liao;Guoliang Chen;Yunquan Zhang

文献摘要

相似文献

为了满足大数据分析的高需求,一些大型分布式集群系统的编程模型被提出并实现,例如MapRe-duce、Dryad和Pregel。然而,与高性能计算领域相比,大数据分析的计算和通信行为的基础和原理还没有得到很好的研究。在本文中,我们回顾了当前的大数据计算模型DOT和DOTA,并提出了一种更通用和实用的模型p-DOT(p-phases DOT)。 p-DOT并不是简单的扩展,而是具有深远的意义:对于一般方面,任何以DOT模型或BSP模型表达的大数据分析作业执行都可以用它来表示;对于实际方面,它会考虑 I/O 行为来评估性能开销。此外,我们提供了一个成本函数,表明对于固定算法和工作负载,最佳机器数量与输入大小的平方根接近线性,并通过多个实验证明了该函数的有效性。
In response to the high demand of big data analytics, several programming models on large and distributed cluster systems have been proposed and implemented, such as MapRe-duce, Dryad and Pregel. However, compared with high performance computing areas, the basis and principles of computation and communication behavior of big data analytics is not well studied. In this paper, we review the current big data computational model DOT and DOTA, and propose a more general and practical model p-DOT (p-phases DOT). p-DOT is not a simple extension, but with profound significance: for general aspects, any big data analytics job execution expressed in DOT model or BSP model can be represented by it; for practical aspects, it considers I/O behavior to evaluate performance overhead. Moreover, we provide a cost function implying that the optimal number of machines is near-linear to the square root of input size for a fixed algorithm and workload, and demonstrate the effectiveness of the function through several experiments.