Dynamic Processing Slots Scheduling for I/O Intensive Jobs of Hadoop MapReduce

Dynamic Processing Slots Scheduling for I/O Intensive Jobs of Hadoop MapReduce
复制标题

DOI:
10.1109/icnc.2012.53
复制
发表时间:
2012-12
期刊:
2012 Third International Conference on Networking and Computing
影响因子:
--
通讯作者:
Shiori Kurazumi;Tomoaki Tsumura;S. Saito;H. Matsuo
Shiori Kurazumi;Tomoaki Tsumura;S. Saito;H. Matsuo
中科院分区:
其他
文献类型:
--
作者:
Shiori Kurazumi;Tomoaki Tsumura;S. Saito;H. Matsuo

文献摘要

相似文献

Hadoop,由Hadoop MapReduce和Hadoop分布式文件系统(HDFS)组成,是用于大规模数据和处理的平台。由于数据的数量一直在全球范围内迅速增加,而且流程的规模变得更大,因此Hadoop吸引了许多云计算企业和技术爱好者,因此分布式处理已变得很普遍。在这种情况下,Hadoop用户正在扩展。我们的研究是为了开发由Hadoop发起的更快的执行工作的速度。在本文中,我们建议使用Hadoop MapReduce的I/O密集作业的动态处理插槽调度,重点关注I/O在执行工作期间等待I/O。当在每个主动任务跟踪器节点上检测到具有高率的I/O等待率的CPU资源时,分配了更多任务以添加免费插槽。我们在Hadoop 1.0.3上实施了我们的方法,在执行时间内提高了大约23%的改善。
Hadoop, consists of Hadoop MapReduce and Hadoop Distributed File System (HDFS), is a platform for large scale data and processing. Distributed processing has become common as the number of data has been increasing rapidly worldwide and the scale of processes has become larger, so that Hadoop has attracted many cloud computing enterprises and technology enthusiasts. Hadoop users are expanding under this situation. Our studies are to develop the faster of executing jobs originated by Hadoop. In this paper, we propose dynamic processing slots scheduling for I/O intensive jobs of Hadoop MapReduce focusing on I/O wait during execution of jobs. Assigning more tasks to added free slots when CPU resources with the high rate of I/O wait have been detected on each active Task Tracker node leads to the improvement of CPU performance. We implemented our method on Hadoop 1.0.3, which results in an improvement of up to about 23% in the execution time.