Building a HighLevel Dataflow System on top of MapReduce: The Pig Experience

Building a HighLevel Dataflow System on top of MapReduce: The Pig Experience
复制标题

DOI:
10.14778/1687553.1687568
复制
发表时间:
2009-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Alan Gates;Olga Natkovich;Shubham Chopra;Pradeep Kamath;Shravan Narayanam;Christopher Olston;B. Reed-B.-Ree
Alan Gates;Olga Natkovich;Shubham Chopra;Pradeep Kamath;Shravan Narayanam;Christopher Olston;B. Reed-B.-Ree
中科院分区:
其他
文献类型:
--
作者:
Alan Gates;Olga Natkovich;Shubham Chopra;Pradeep Kamath;Shravan Narayanam;Christopher Olston;B. Reed-B.-Ree

文献摘要

被引文献

相似文献

越来越多的组织捕获、转换和分析庞大的数据集。突出的例子包括互联网公司和电子科学。Map-Reduce可伸缩的XML范式已经成为这些应用程序的流行。它简单、显式的XML编程模型比传统的高级声明式方法SQL更受欢迎。另一方面,Map-Reduce的极端简单性导致了许多低级的黑客来处理实践中出现的多步骤、分支的问题。此外,用户必须重复编写标准操作,如手工连接。这些实践浪费时间,引入错误,损害可读性,并阻碍优化。Pig是一个高级的XML系统,旨在SQL和Map-Reduce之间的最佳点。Pig提供了SQL风格的高级数据操作结构,可以在显式的XML中组装,并与自定义Map和Reduce风格的函数或可执行文件交叉使用。Pig程序被编译成Map-Reduce作业序列,并在Hadoop Map-Reduce环境中执行。Pig和Hadoop都是由Apache软件基金会管理的开源项目。本文描述了我们在开发Pig时所面临的挑战,并报告了Pig执行和原始Map-Reduce执行之间的性能比较。
Increasingly, organizations capture, transform and analyze enormous data sets. Prominent examples include internet companies and e-science. The Map-Reduce scalable dataflow paradigm has become popular for these applications. Its simple, explicit dataflow programming model is favored by some over the traditional high-level declarative approach: SQL. On the other hand, the extreme simplicity of Map-Reduce leads to much low-level hacking to deal with the many-step, branching dataflows that arise in practice. Moreover, users must repeatedly code standard operations such as join by hand. These practices waste time, introduce bugs, harm readability, and impede optimizations. Pig is a high-level dataflow system that aims at a sweet spot between SQL and Map-Reduce. Pig offers SQL-style high-level data manipulation constructs, which can be assembled in an explicit dataflow and interleaved with custom Map- and Reduce-style functions or executables. Pig programs are compiled into sequences of Map-Reduce jobs, and executed in the Hadoop Map-Reduce environment. Both Pig and Hadoop are open-source projects administered by the Apache Software Foundation. This paper describes the challenges we faced in developing Pig, and reports performance comparisons between Pig execution and raw Map-Reduce execution.