Handling partitioning skew in MapReduce using LEEN

Handling partitioning skew in MapReduce using LEEN
复制标题

使用 LEEN 处理 MapReduce 中的分区偏差

DOI:
10.1007/s12083-013-0213-7
复制
发表时间:
2013-05
影响因子:
4.2
通讯作者:
Song Wu
Song Wu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shadi Ibrahim;Hai Jin;Lu Lu;Bingsheng He;Gabriel Antoniu;Song Wu

文献摘要

被引文献

相似文献

MapReduce正在成为一个重要的大数据处理工具。数据局域性是MapReduce的一个关键特性,在数据密集型云系统中被广泛利用:它通过共同分配计算和数据存储来避免在处理大量数据时网络饱和,特别是在地图阶段。然而,我们对广泛使用的MapReduce实现Hadoop的研究表明,分区倾斜(分区倾斜指的是中间键的频率或分布的变化,或不同数据节点之间两者的变化)的存在会导致shuffle阶段的大量数据传输,并导致不同数据节点之间reduce输入的显著不公平。结果,由于shuffle阶段的长时间数据传输以及计算倾斜,特别是在reduce阶段,应用程序的性能严重下降。在本文中,我们开发了一种新的算法LEEN,用于MapReduce中位置感知和公平感知的键划分。LEEN采用异步映射和缩减方案。所有缓冲的中间键都根据它们的频率和洗牌阶段后预期数据分布的公平性进行分区。我们已经将LEEN集成到Hadoop中。实验表明,LEEN可以有效地实现更高的局部性,并减少洗牌数据的数量。更重要的是,LEEN保证了减少投入的公平分配。因此,LEEN在不同的工作负载上实现了高达45%的性能改进。
MapReduce is emerging as a prominent tool for big data processing. Data locality is a key feature in MapReduce that is extensively leveraged in data-intensive cloud systems: it avoids network saturation when processing large amounts of data by co-allocating computation and data storage, particularly for the map phase. However, our studies with Hadoop, a widely used MapReduce implementation, demonstrate that the presence of partitioning skew (Partitioning skew refers to the case when a variation in either the intermediate keys’ frequencies or their distributions or both among different data nodes) causes a huge amount of data transfer during the shuffle phase and leads to significant unfairness on the reduce input among different data nodes. As a result, the applications severe performance degradation due to the long data transfer during the shuffle phase along with the computation skew, particularly in reduce phase. In this paper, we develop a novel algorithm named LEEN for locality-aware and fairness-aware key partitioning in MapReduce. LEEN embraces an asynchronous map and reduce scheme. All buffered intermediate keys are partitioned according to their frequencies and the fairness of the expected data distribution after the shuffle phase. We have integrated LEEN into Hadoop. Our experiments demonstrate that LEEN can efficiently achieve higher locality and reduce the amount of shuffled data. More importantly, LEEN guarantees fair distribution of the reduce inputs. As a result, LEEN achieves a performance improvement of up to 45 % on different workloads.