Improving MapReduce Performance with Partial Speculative Execution

Improving MapReduce Performance with Partial Speculative Execution
复制标题

DOI:
10.1007/s10723-015-9350-y
复制
发表时间:
2015-12
影响因子:
5.5
通讯作者:
Yaoguang Wang;Weiming Lu;Renjie Lou;Baogang Wei
Yaoguang Wang;Weiming Lu;Renjie Lou;Baogang Wei
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yaoguang Wang;Weiming Lu;Renjie Lou;Baogang Wei

文献摘要

被引文献

相似文献

The MapReduce framework has become the de facto standard for big data processing due to its attractive features and abilities. One is that it automatically parallelizes a job into multiple tasks and transparently handles task execution on a large cluster of commodity machines. The increasing heterogeneity of distributed environments may result in a few straggling tasks, which prolong job completion. Speculative execution is proposed to mitigate stragglers. However, the existing speculative execution mechanism could not work efficiently as many speculative tasks are still slower than their original tasks. In this paper, we explore an approach to increase the efficiency of speculative execution, and further improve MapReduce performance. We propose thePartialSpeculativeExecution (PSE) strategy to make speculative tasks start from the checkpoint. By leveraging the checkpoint of original tasks, PSE can eliminate the costs of re-reading, re-copying, and re-computing the processed data. We implement PSE in Hadoop, and evaluate its performance in terms of job completion time and the efficiency of speculative execution under several kinds of classical workloads. Experimental results show that, in heterogeneous environments with stragglers, PSE completes jobs 56 % faster than that with no speculation and 12 % faster than that with LATE, an improved speculative execution algorithm. In addition, on average PSE can improve the efficiency of speculative execution by 24 % compared to LATE.