Continuously Improving the Resource Utilization of Iterative Parallel Dataflows

Continuously Improving the Resource Utilization of Iterative Parallel Dataflows
复制标题

不断提高迭代并行数据流的资源利用率

DOI:
10.1109/icdcsw.2016.20
复制
发表时间:
2016
期刊:
2016 IEEE 36th International Conference on Distributed Computing Systems Workshops (ICDCSW)
影响因子:
--
通讯作者:
O. Kao
O. Kao
中科院分区:
--
文献类型:
--
作者:
L. Thamsen;T. Renner;O. Kao

文献摘要

参考文献

被引文献

相似文献

Apache Flink 等并行数据流系统允许使用迭代程序分析大型数据集。然而,为此类作业分配一组经济有效的资源是一项艰巨的任务,因为资源利用率取决于许多因素,例如数据集大小、键值分布、程序的计算复杂性和底层硬件。更重要的是,其中一些因素在执行之前并不为人所知。例如,通常事先没有可用的数据统计信息,例如键值分布。因此,我们建议利用迭代数据流程序的重复性来提高运行时的资源利用率。根据先前迭代中收集的运行时统计数据,在迭代之间的同步障碍处动态调整资源分配。这种方法有两个优点:首先,即使对于并行执行的任务管道,在障碍处也可以获得详细的统计数据。其次,在障碍处,可以调整数据流,而无需复杂地处理中间任务状态。本文提出了与 Apache Flink 集成的原型以及对 480 核集群的评估。一项实验显示,通过在更短的时间内分配更多资源,作业运行时间减少了 57%,另一项实验则在不显着延长作业运行时间的情况下释放了高达 40% 的剩余资源。
Parallel dataflow systems like Apache Flink allow analysis of large datasets with iterative programs. However, allocating a cost-effective set of resources for such jobs is a difficult task as the resource utilization depends on many factors such as dataset size, key value distributions, computational complexity of programs, and the underlying hardware. What's more, some of these factors are not well known before the execution. There are, for example, often no data statistics such as key value distributions available beforehand. For this reason, we propose to improve the resource utilization at runtime using the repetitive nature of iterative dataflow programs. Based on runtime statistics gathered in previous iterations, the resource allocation is adapted dynamically at the synchronization barriers between iterations. This approach has two advantages: First, at barriers detailed statistics can be available, even for parallelly executed task pipelines. Second, at barriers dataflows can be adapted without complex handling of intermediate task state. This paper presents a prototype integrated with Apache Flink and an evaluation on a cluster with 480 cores. One experiment shows a 57% reduction of the job runtime by allocating more resources for a shorter time, another experiment a release of up to 40% surplus resources without significantly extending the job runtime.
检测基于 DAG 的并行数据流程序中的瓶颈
DOI: 10.1109/mtags.2010.5699429
发表时间: 2010
期刊: 2010 3rd Workshop on Many-Task Computing on Grids and Supercomputers
影响因子: --
作者:
Dominic Battré;Matthias Hovestadt;Björn Lohrmann;Alexander Stanik;Daniel Warneke
通讯作者: Daniel Warneke
DOI: 10.1007/s00778-014-0357-y
发表时间: 2014-12-01
期刊: VLDB JOURNAL
影响因子: 4.2
作者:
Alexandrov, Alexander;Bergmann, Rico;Warneke, Daniel
通讯作者: Warneke, Daniel
DOI: 10.1145/2463676.2465282
发表时间: 2013-06
期刊: --
影响因子: --
作者:
R. Fernandez;Matteo Migliavacca;Evangelia Kalyvianaki;P. Pietzuch
通讯作者: R. Fernandez;Matteo Migliavacca;Evangelia Kalyvianaki;P. Pietzuch
DOI: 10.1109/icdcs.2015.48
发表时间: 2015-06
期刊: 2015 IEEE 35th International Conference on Distributed Computing Systems
影响因子: --
作者:
Björn Lohrmann;P. Janacik;O. Kao
通讯作者: Björn Lohrmann;P. Janacik;O. Kao
DOI: 10.14778/2350229.2350245
发表时间: 2012-07
期刊: Proc. VLDB Endow.
影响因子: --
作者:
Stephan Ewen;K. Tzoumas;Moritz Kaufmann;V. Markl
通讯作者: Stephan Ewen;K. Tzoumas;Moritz Kaufmann;V. Markl