Can High-Performance Interconnects Benefit Hadoop Distributed File System ?
Can High-Performance Interconnects Benefit Hadoop Distributed File System ?
复制标题
DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
S. Sur;Hao Wang;Jian Huang;Xiangyong Ouyang;D. Panda
中科院分区:
文献类型:
--
作者:
S. Sur;Hao Wang;Jian Huang;Xiangyong Ouyang;D. Panda
During the past several years, the MapReduce computing model has emerged as a scalable model that is capable of processing petabytes of data. The Hadoop MapReduce framework, has enabled large scale Internet applications and has been adopted by many organizations. The Hadoop Distributed File System (HDFS) lies at the heart of the ecosystem of software. It was designed to operate and scale on commodity hardware such as cheap Linux machines connected with Gigabit Ethernet. The field of Highperformance Computing (HPC) has been witnessing a transition to commodity clusters. Increasingly, networks such as InfiniBand and 10Gigabit Ethernet have become commoditized and available on motherboards and as lowcost PCI-Express devices. Software that drives InfiniBand and 10Gigabit Ethernet works on mainline Linux kernels. These interconnects provide high bandwidth along with low CPU utilization for network intensive operations. As the amounts of data processed by Internet applications reaches hundreds of petabytes, it is expected that the network performance will be a key component towards scaling data-centers. In this paper, we examine the impact of high-performance interconnects on HDFS. Our findings reveal that the impact is substantial. We observe up to 11%, 30% and 100% performance improvement for the sort, random write and sequential write benchmarks using magnetic disks (HDD). We also find that with the emerging trend of Solid State Drives, having a faster interconnect makes a larger impact as local I/O costs are reduced. We observe up to 48%, 59% and 219% improvement for the same benchmarks when SSD is used in combination with advanced interconnection networks and protocols.