Scaling up MapReduce-based Big Data Processing on Multi-GPU systems

Scaling up MapReduce-based Big Data Processing on Multi-GPU systems
复制标题

DOI:
10.1007/s10586-014-0400-1
复制
发表时间:
2015-03
期刊:
Cluster Computing
影响因子:
--
通讯作者:
Hai Jiang;Yi Chen;Zhi Qiao;Tien-Hsiung Weng;Kuan-Ching Li
Hai Jiang;Yi Chen;Zhi Qiao;Tien-Hsiung Weng;Kuan-Ching Li
中科院分区:
其他
文献类型:
--
作者:
Hai Jiang;Yi Chen;Zhi Qiao;Tien-Hsiung Weng;Kuan-Ching Li

文献摘要

被引文献

相似文献

MapReduce是一种流行的数据并行处理模型,包含了计算技术的最新进展,并已被广泛用于大规模数据分析。对MapReduce的高需求刺激了对不同架构模型和计算范例的MapReduce实现的研究,如多核集群、云、立体板和GPU。特别是,现有的基于GPU的MapReduce方法主要集中在单GPU算法上,由于GPU内存容量有限,不能处理大数据集。本文在原有的多GPU MapReduceMGMR版本的基础上,提出了一种消除GPU内存限制的升级版本MGMR++和一种流水线版本PMGMR,以通过CPU内存和硬盘来应对大数据挑战。MGMR++是从MGMR扩展而来的,具有灵活的C++模板和CPU内存利用率,而PMGMR通过流和Hyper-Q等最新的GPU功能以及硬盘利用率对性能进行了微调。与MGMR(酱等人,集群计算2013)相比,所提出的方案实现了约2.5倍的性能提升,增加了系统可伸缩性,并允许程序员为大数据编写简单的MapReduce代码。
MapReduce is a popular data-parallel processing model encompassed with recent advances in computing technology and has been widely exploited for large-scale data analysis. The high demand on MapReduce has stimulated the investigation of MapReduce implementations with different architectural models and computing paradigms, such as multi-core clusters, Clouds, Cubieboards and GPUs. Particularly, current GPU-based MapReduce approaches mainly focus on single-GPU algorithms and cannot handle large data sets, due to the limited GPU memory capacity. Based on the previous multi-GPU MapReduce version MGMR, this paper proposes an upgrade version MGMR++ to eliminate GPU memory limitation and a pipelined version, PMGMR, to handle the Big Data challenge through both CPU memory and hard disks. MGMR++ is extended from MGMR with flexible C++ templates and CPU memory utilization, while PMGMR fine-tuned the performance through the latest GPU features such as streams and Hyper-Q as well as hard disk utilization. Compared to MGMR (Jiang et al., Cluster Computing 2013), the proposed schemes achieve about 2.5-fold performance improvement, increase system scalability, and allow programmers to write straightforward MapReduce code for Big Data.