MapIterativeReduce: a framework for reduction-intensive data processing on azure clouds
MapIterativeReduce: a framework for reduction-intensive data processing on azure clouds
复制标题
DOI:
10.1145/2287016.2287019
复制
发表时间:
2012-06
期刊:
影响因子:
--
通讯作者:
R. Tudoran;Alexandru Costan;Gabriel Antoniu
中科院分区:
文献类型:
--
作者:
R. Tudoran;Alexandru Costan;Gabriel Antoniu
With the emergence of cloud computing as an alternative to supercomputers to support data intensive applications, MapReduce has arisen as a major programming model for data analysis on clouds. In this context, reduce-intensive algorithms are becoming increasingly useful in applications such as data clustering, classification and mining. However, platforms like MapReduce or Dryad lack built-in support for reduce-intensive workloads. This paper introduces MapIterativeReduce, a framework which 1) extends the MapReduce programming model to better support reduce-intensive applications and 2) substantially improves their efficiency by eliminating the implicit barrier between the Map and the Reduce phase. We evaluated MapIterativeReduce on the Microsoft Azure cloud with synthetic benchmarks and with a real-life application. Compared to state-of-art solutions, our approach reduces the execution times by up to 75%.