Can High-Performance Interconnects Benefit Hadoop Distributed File System ?

Can High-Performance Interconnects Benefit Hadoop Distributed File System ?
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
S. Sur;Hao Wang;Jian Huang;Xiangyong Ouyang;D. Panda
S. Sur;Hao Wang;Jian Huang;Xiangyong Ouyang;D. Panda
中科院分区:
其他
文献类型:
--
作者:
S. Sur;Hao Wang;Jian Huang;Xiangyong Ouyang;D. Panda

文献摘要

被引文献

相似文献

在过去的几年中,MapReduce计算模型已经成为一种能够处理PB级数据的可扩展模型。Hadoop MapReduce框架已经实现了大规模的互联网应用程序,并已被许多组织采用。Hadoop分布式文件系统(HDFS)是软件生态系统的核心。它被设计为在商用硬件上运行和扩展,例如与千兆以太网连接的廉价Linux机器。高性能计算(HPC)领域一直在向商品集群过渡。InfiniBand和10千兆以太网等网络已经越来越商品化,可以在主板上使用,也可以作为低成本的PCI-Express设备使用。驱动InfiniBand和10千兆以太网的软件在主线Linux内核上工作。这些互连为网络密集型操作提供高带宽沿着低CPU利用率。随着互联网应用程序处理的数据量达到数百PB,预计网络性能将成为扩展数据中心的关键组成部分。在本文中,我们研究了高性能互连对HDFS的影响。我们的研究结果表明,影响是巨大的。我们观察到高达11%,30%和100%的性能提高排序,随机写入和顺序写入基准使用磁盘(HDD)。我们还发现,随着固态硬盘的新兴趋势,随着本地I/O成本的降低,更快的互连会产生更大的影响。当SSD与先进的互连网络和协议结合使用时,我们观察到相同基准的改进高达48%,59%和219%。
During the past several years, the MapReduce computing model has emerged as a scalable model that is capable of processing petabytes of data. The Hadoop MapReduce framework, has enabled large scale Internet applications and has been adopted by many organizations. The Hadoop Distributed File System (HDFS) lies at the heart of the ecosystem of software. It was designed to operate and scale on commodity hardware such as cheap Linux machines connected with Gigabit Ethernet. The field of Highperformance Computing (HPC) has been witnessing a transition to commodity clusters. Increasingly, networks such as InfiniBand and 10Gigabit Ethernet have become commoditized and available on motherboards and as lowcost PCI-Express devices. Software that drives InfiniBand and 10Gigabit Ethernet works on mainline Linux kernels. These interconnects provide high bandwidth along with low CPU utilization for network intensive operations. As the amounts of data processed by Internet applications reaches hundreds of petabytes, it is expected that the network performance will be a key component towards scaling data-centers. In this paper, we examine the impact of high-performance interconnects on HDFS. Our findings reveal that the impact is substantial. We observe up to 11%, 30% and 100% performance improvement for the sort, random write and sequential write benchmarks using magnetic disks (HDD). We also find that with the emerging trend of Solid State Drives, having a faster interconnect makes a larger impact as local I/O costs are reduced. We observe up to 48%, 59% and 219% improvement for the same benchmarks when SSD is used in combination with advanced interconnection networks and protocols.