Abstract cost models for distributed data-intensive computations
Abstract cost models for distributed data-intensive computations
复制标题
分布式数据密集型计算的抽象成本模型
DOI:
10.1007/s10619-018-7244-2
复制
发表时间:
2019
影响因子:
1.2
通讯作者:
Yao, Yi
中科院分区:
文献类型:
--
作者:
Li, Rundong;Mi, Ningfang;Riedewald, Mirek;Sun, Yizhou;Yao, Yi
We consider data analytics workloads on distributed architectures, in particular clusters of commodity machines. To find a job partitioning that minimizes running time, a cost model, which we more accurately refer to as makespan model, is needed. In attempting to find the simplest possible, but sufficiently accurate, such model, we explore piecewise linear functions of input, output, and computational complexity. They are abstract in the sense that they capture fundamental algorithm properties, but do not require explicit modeling of system and implementation details such as the number of disk accesses. We show how the simplified functional structure can be exploited to reduce optimization cost. In the general case, we identify a lower bound that can be used for search-space pruning. For applications with homogeneous tasks, we further demonstrate how to directly integrate the model into the makespan optimization process, reducing search-space dimensionality and thus complexity by orders of magnitude. Experimental results provide evidence of good prediction quality and successful makespan optimization across a variety of operators and cluster architectures.
DOI:
10.14778/2733004.2733005
发表时间:
2014-08
期刊:
Proc. VLDB Endow.
影响因子:
--
作者:
Juwei Shi;Jia Zou;Jiaheng Lu;Zhao Cao;Shiqiang Li;Chen Wang
通讯作者:
Juwei Shi;Jia Zou;Jiaheng Lu;Zhao Cao;Shiqiang Li;Chen Wang
影响因子:
22.7
作者:
R. Arkin
通讯作者:
R. Arkin