A Data Locality and Skew Aware Task Scheduler for MapReduce in Cloud Computing
A Data Locality and Skew Aware Task Scheduler for MapReduce in Cloud Computing
复制标题
DOI:
10.5339/qfarf.2011.csp21
复制
发表时间:
2011-12
期刊:
影响因子:
--
通讯作者:
Mohammad Hammoud;Suhail Rehman;M. Sakr
中科院分区:
文献类型:
--
作者:
Mohammad Hammoud;Suhail Rehman;M. Sakr
Abstract Inspired by the success and the increasing prevalence of MapReduce, this work proposes a novel MapReduce task scheduler. MapReduce is by far one of the most successful realizations of large-scale, data-intensive, cloud computing platforms. As compared to traditional programming models, MapReduce automatically and efficiently parallelizes computation by running multiple Map and/or Reduce tasks over distributed data across multiple machines. Hadoop, an open source implementation of MapReduce, schedules Map tasks in the vicinity of their input splits seeking diminished network traffic. However, when Hadoop schedules Reduce tasks, it neither exploits data locality nor addresses data partitioning skew inherent in many MapReduce applications. Consequently, MapReduce experiences a performance penalty and network congestion as observed in our experimental results. Recently there has been some work concerned with leveraging data locality in Reduce task scheduling. For instance, one study suggests a locali...