Towards Systematic Parallel Programming over MapReduce

Towards Systematic Parallel Programming over MapReduce
复制标题

DOI:
10.1007/978-3-642-23397-5_5
复制
发表时间:
2011-08
期刊:
--
影响因子:
--
通讯作者:
Yu Liu;Zhenjiang Hu;Kiminori Matsuzaki
Yu Liu;Zhenjiang Hu;Kiminori Matsuzaki
中科院分区:
其他
文献类型:
--
作者:
Yu Liu;Zhenjiang Hu;Kiminori Matsuzaki

文献摘要

相似文献

MapReduce是一种适用于数据密集型分布式并行计算的实用而流行的编程模型。但是,用MapReduce语言系统地开发并行程序仍然是一个挑战,因为要得到一个与MapReduce语言匹配的合适的分而治之的算法通常并不容易。本文利用列表同态的程序计算理论,提出了一种基于同态的系统并行编程框架ScrewDriver。螺丝刀被实现为Hadoop之上的Java库。对于满足第三同态定理要求的两个顺序函数可以解决的任何问题,ScrewDriver可以自动推导出列表同态的并行算法,并将初始的顺序程序转换为高效的MapReduce程序。用户既不需要关心并行性,也不需要对MapReduce有深入的了解。除了我们框架的编程模型的简单性之外,这种计算方法还使我们能够解决许多问题,而使用MapReduce直接解决这些问题并不是轻而易举的事情。
MapReduce is a useful and popular programming model for data-intensive distributed parallel computing. But it is still a challenge to develop parallel programs with MapReduce systematically, since it is usually not easy to derive a proper divide-and-conquer algorithm that matches MapReduce. In this paper, we propose a homomorphism-based framework named Screwdriver for systematic parallel programming with MapReduce, making use of the program calculation theory of list homomorphisms. Screwdriver is implemented as a Java library on top of Hadoop. For any problem which can be resolved by two sequential functions that satisfy the requirements of the third homomorphism theorem, Screwdriver can automatically derive a parallel algorithm as a list homomorphism and transform the initial sequential programs to an efficient MapReduce program. Users need neither to care about parallelism nor to have deep knowledge of MapReduce. In addition to the simplicity of the programming model of our framework, such a calculational approach enables us to resolve many problems that it would be nontrivial to resolve directly with MapReduce.