A comparison of join algorithms for log processing in MaPreduce

A comparison of join algorithms for log processing in MaPreduce
复制标题

DOI:
10.1145/1807167.1807273
复制
发表时间:
2010-06
期刊:
Proceedings of the 2010 ACM SIGMOD International Conference on Management of data
影响因子:
--
通讯作者:
Spyros Blanas;J. Patel;V. Ercegovac;Jun Rao;E. Shekita;Yuanyuan Tian
Spyros Blanas;J. Patel;V. Ercegovac;Jun Rao;E. Shekita;Yuanyuan Tian
中科院分区:
其他
文献类型:
--
作者:
Spyros Blanas;J. Patel;V. Ercegovac;Jun Rao;E. Shekita;Yuanyuan Tian

文献摘要

被引文献

相似文献

MapReduce框架越来越多地被用于分析大量数据。使用MapReduce完成的一种重要的数据分析类型是日志处理,其中对点击流或事件日志进行过滤、聚合或挖掘模式。作为分析的一部分,通常需要将日志与诸如用户信息之类的参考数据结合起来。虽然已经有很多研究在并行和分布式dbms中检查连接算法,但是MapReduce框架对于连接来说是很麻烦的。MapReduce程序员经常使用简单但效率低下的算法来执行连接。在本文中,我们描述了MapReduce中一些众所周知的连接策略的关键实现细节,并在100节点Hadoop集群上对这些连接技术进行了全面的实验比较。我们的结果为MapReduce平台提供了独特的见解,并提供了在该平台上何时使用特定连接算法的指导。
The MapReduce framework is increasingly being used to analyze large volumes of data. One important type of data analysis done with MapReduce is log processing, in which a click-stream or an event log is filtered, aggregated, or mined for patterns. As part of this analysis, the log often needs to be joined with reference data such as information about users. Although there have been many studies examining join algorithms in parallel and distributed DBMSs, the MapReduce framework is cumbersome for joins. MapReduce programmers often use simple but inefficient algorithms to perform joins. In this paper, we describe crucial implementation details of a number of well-known join strategies in MapReduce, and present a comprehensive experimental comparison of these join techniques on a 100-node Hadoop cluster. Our results provide insights that are unique to the MapReduce platform and offer guidance on when to use a particular join algorithm on this platform.