Hadoop++

Hadoop++
复制标题

DOI:
10.14778/1920841.1920908
复制
发表时间:
2010-09
影响因子:
2.5
通讯作者:
J. Dittrich;Jorge-Arnulfo Quiané-Ruiz;Alekh Jindal;Y. Kargin;Vinay Setty;Jörg Schad
J. Dittrich;Jorge-Arnulfo Quiané-Ruiz;Alekh Jindal;Y. Kargin;Vinay Setty;Jörg Schad
中科院分区:
计算机科学2区
文献类型:
--
作者:
J. Dittrich;Jorge-Arnulfo Quiané-Ruiz;Alekh Jindal;Y. Kargin;Vinay Setty;Jörg Schad

文献摘要

被引文献

相似文献

MapReduce是一种计算范式,近年来在工业界和研究领域都受到了广泛关注。与并行数据库管理系统不同,MapReduce允许非专业用户在非常大的集群和云端的超大型数据集上运行复杂的分析任务。然而,这是有代价的:MapReduce以面向扫描的方式处理任务。因此,Hadoop(MapReduce的一种开源实现)的性能往往比不上配置良好的并行数据库管理系统。在本文中,我们提出了一种名为Hadoop++的新型系统:它在完全不改变Hadoop框架的情况下提高任务性能(Hadoop甚至都“察觉不到”)。为了实现这一目标,我们不是去改变一个正在运行的系统(Hadoop),而是仅通过用户定义函数(UDFs)在合适的位置注入我们的技术,从内部影响Hadoop。这带来了三个重要的结果:首先,Hadoop++的性能明显优于Hadoop。其次,Hadoop未来的任何变更都可以直接在Hadoop++中使用,而无需重写任何粘合代码。第三,Hadoop++不需要改变Hadoop的接口。我们的实验表明,在与索引和连接处理相关的任务方面,Hadoop++优于Hadoop和HadoopDB。
MapReduce is a computing paradigm that has gained a lot of attention in recent years from industry and research. Unlike parallel DBMSs, MapReduce allows non-expert users to run complex analytical tasks over very large data sets on very large clusters and clouds. However, this comes at a price: MapReduce processes tasks in a scan-oriented fashion. Hence, the performance of Hadoop --- an open-source implementation of MapReduce --- often does not match the one of a well-configured parallel DBMS. In this paper we propose a new type of system named Hadoop++: it boosts task performance without changing the Hadoop framework at all (Hadoop does not even 'notice it'). To reach this goal, rather than changing a working system (Hadoop), we inject our technology at the right places through UDFs only and affect Hadoop from inside. This has three important consequences: First, Hadoop++ significantly outperforms Hadoop. Second, any future changes of Hadoop may directly be used with Hadoop++ without rewriting any glue code. Third, Hadoop++ does not need to change the Hadoop interface. Our experiments show the superiority of Hadoop++ over both Hadoop and HadoopDB for tasks related to indexing and join processing.