MapIterativeReduce: a framework for reduction-intensive data processing on azure clouds

MapIterativeReduce: a framework for reduction-intensive data processing on azure clouds
复制标题

DOI:
10.1145/2287016.2287019
复制
发表时间:
2012-06
期刊:
--
影响因子:
--
通讯作者:
R. Tudoran;Alexandru Costan;Gabriel Antoniu
R. Tudoran;Alexandru Costan;Gabriel Antoniu
中科院分区:
其他
文献类型:
--
作者:
R. Tudoran;Alexandru Costan;Gabriel Antoniu

文献摘要

被引文献

相似文献

随着云计算作为超级计算机的替代品来支持数据密集型应用程序的出现,MapReduce已经成为云数据分析的主要编程模型。在这种情况下,减少密集型算法在数据聚类,分类和挖掘等应用中变得越来越有用。然而,像MapReduce或Dryad这样的平台缺乏对减少密集型工作负载的内置支持。本文介绍了MapIterativeReduce,一个框架,1)扩展MapReduce编程模型,以更好地支持减少密集型应用程序,2)通过消除Map和Reduce阶段之间的隐式障碍,大大提高了它们的效率。我们使用合成基准和实际应用程序在Microsoft Azure云上评估了MapIterativeReduce。与最先进的解决方案相比,我们的方法将执行时间减少了75%。
With the emergence of cloud computing as an alternative to supercomputers to support data intensive applications, MapReduce has arisen as a major programming model for data analysis on clouds. In this context, reduce-intensive algorithms are becoming increasingly useful in applications such as data clustering, classification and mining. However, platforms like MapReduce or Dryad lack built-in support for reduce-intensive workloads. This paper introduces MapIterativeReduce, a framework which 1) extends the MapReduce programming model to better support reduce-intensive applications and 2) substantially improves their efficiency by eliminating the implicit barrier between the Map and the Reduce phase. We evaluated MapIterativeReduce on the Microsoft Azure cloud with synthetic benchmarks and with a real-life application. Compared to state-of-art solutions, our approach reduces the execution times by up to 75%.