Distributed Programming with MapReduce

Distributed Programming with MapReduce
复制标题

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
J. Dean;Sanjay Ghemawat
J. Dean;Sanjay Ghemawat
中科院分区:
其他
文献类型:
--
作者:
J. Dean;Sanjay Ghemawat

文献摘要

被引文献

相似文献

这一章描述了MAPREDUCE的设计和实现,MAPREDUCE是一个用于大规模数据处理问题的编程系统。MapReduce是作为一种简化大规模计算开发的方法而开发的。MapReduce程序被自动并行化,并在一个大型的商用机器集群上执行。运行时系统负责对输入数据进行分区、在一组机器上调度程序的执行、处理机器故障以及管理所需的机器间通信等细节。这使得没有任何并行和分布式系统经验的程序员可以轻松地利用大型分布式系统的资源。
THIS CHAPTER DESCRIBES THE DESIGN AND IMPLEMENTATION OF MAPREDUCE, a programming system for large-scale data processing problems. MapReduce was developed as a way of simplifying the development of large-scale computations at Google. MapReduce programs are automatically parallelized and executed on a large cluster of commodity machines. The runtime system takes care of the details of partitioning the input data, scheduling the program’s execution across a set of machines, handling machine failures, and managing the required intermachine communication. This allows programmers without any experience with parallel and distributed systems to easily utilize the resources of a large distributed system.