A Data Locality and Skew Aware Task Scheduler for MapReduce in Cloud Computing

A Data Locality and Skew Aware Task Scheduler for MapReduce in Cloud Computing
复制标题

DOI:
10.5339/qfarf.2011.csp21
复制
发表时间:
2011-12
期刊:
--
影响因子:
--
通讯作者:
Mohammad Hammoud;Suhail Rehman;M. Sakr
Mohammad Hammoud;Suhail Rehman;M. Sakr
中科院分区:
其他
文献类型:
--
作者:
Mohammad Hammoud;Suhail Rehman;M. Sakr

文献摘要

被引文献

相似文献

摘要受MapReduce的成功和日益流行的启发,本文提出了一种新的MapReduce任务调度器。MapReduce是迄今为止大规模、数据密集型云计算平台最成功的实现之一。与传统的编程模型相比,MapReduce通过在多台机器上的分布式数据上运行多个Map和/或Reduce任务来自动有效地并行计算。Hadoop是MapReduce的一个开源实现,它将Map任务安排在输入拆分附近,以减少网络流量。然而,当Hadoop调度Reduce任务时,它既没有利用数据局部性,也没有解决许多MapReduce应用程序中固有的数据分区偏斜。因此,正如我们的实验结果所观察到的那样,MapReduce经历了性能损失和网络拥堵。最近有一些工作涉及在Reduce任务调度中利用数据局部性。例如,一项研究表明,当地...
Abstract Inspired by the success and the increasing prevalence of MapReduce, this work proposes a novel MapReduce task scheduler. MapReduce is by far one of the most successful realizations of large-scale, data-intensive, cloud computing platforms. As compared to traditional programming models, MapReduce automatically and efficiently parallelizes computation by running multiple Map and/or Reduce tasks over distributed data across multiple machines. Hadoop, an open source implementation of MapReduce, schedules Map tasks in the vicinity of their input splits seeking diminished network traffic. However, when Hadoop schedules Reduce tasks, it neither exploits data locality nor addresses data partitioning skew inherent in many MapReduce applications. Consequently, MapReduce experiences a performance penalty and network congestion as observed in our experimental results. Recently there has been some work concerned with leveraging data locality in Reduce task scheduling. For instance, one study suggests a locali...