Only Aggressive Elephants are Fast Elephants

Only Aggressive Elephants are Fast Elephants
复制标题

DOI:
10.14778/2350229.2350272
复制
发表时间:
2012-07
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Dittrich;Jorge-Arnulfo Quiané-Ruiz;Stefan Richter;Stefan Schuh;Alekh Jindal;Jörg Schad
J. Dittrich;Jorge-Arnulfo Quiané-Ruiz;Stefan Richter;Stefan Schuh;Alekh Jindal;Jörg Schad
中科院分区:
其他
文献类型:
--
作者:
J. Dittrich;Jorge-Arnulfo Quiané-Ruiz;Stefan Richter;Stefan Schuh;Alekh Jindal;Jörg Schad

文献摘要

被引文献

相似文献

黄色的大象很慢。一个主要原因是,它们在响应大象骑手的命令之前就完全消耗了自己的输入。一些聪明的骑手已经训练他们的黄象在响应之前只消耗部分输入。然而,教大象做这件事的时间很长。如此之高,教学课程往往没有回报。我们采取不同的方法。我们让大象具有攻击性;只有这样才能使它们跑得很快。我们提出了HAIL(Hadoop积极索引库),HDFS和Hadoop MapReduce的增强,大大提高了几类MapReduce作业的运行时间。HAIL改变HDFS的上传管道,以便在每个数据块副本上创建不同的聚集索引。HAIL的一个有趣的特性是,我们通常会创造一个双赢的局面:我们改进了HDFS的数据上传和实际Hadoop MapReduce作业的运行时。在数据上传方面,HAIL比HDFS提高了60%,默认复制因子为3。在查询执行方面,我们证明了HAIL的运行速度比Hadoop快68倍。在我们的实验中,我们使用六个集群,包括物理和EC2集群多达100个节点。一系列的可扩展性实验也证明了HAIL的优越性。
Yellow elephants are slow. A major reason is that they consume their inputs entirely before responding to an elephant rider's orders. Some clever riders have trained their yellow elephants to only consume parts of the inputs before responding. However, the teaching time to make an elephant do that is high. So high that the teaching lessons often do not pay off. We take a different approach. We make elephants aggressive; only this will make them very fast. We propose HAIL (Hadoop Aggressive Indexing Library), an enhancement of HDFS and Hadoop MapReduce that dramatically improves runtimes of several classes of MapReduce jobs. HAIL changes the upload pipeline of HDFS in order to create different clustered indexes on each data block replica. An interesting feature of HAIL is that we typically create a win-win situation: we improve both data upload to HDFS and the runtime of the actual Hadoop MapReduce job. In terms of data upload, HAIL improves over HDFS by up to 60% with the default replication factor of three. In terms of query execution, we demonstrate that HAIL runs up to 68x faster than Hadoop. In our experiments, we use six clusters including physical and EC2 clusters of up to 100 nodes. A series of scalability experiments also demonstrates the superiority of HAIL.