Exploiting Dynamic Resource Allocation for Efficient Parallel Data Processing in the Cloud

Exploiting Dynamic Resource Allocation for Efficient Parallel Data Processing in the Cloud
复制标题

DOI:
10.1109/tpds.2011.65
复制
发表时间:
2011-06-01
影响因子:
5.3
通讯作者:
Kao, Odej
Kao, Odej
中科院分区:
计算机科学2区
文献类型:
--
作者:
Warneke, Daniel;Kao, Odej

文献摘要

被引文献

相似文献

近年来,临时数据处理已成为基础架构-AS-A-Service(IaaS)云的杀手应用之一。主要的云计算公司已开始在其产品组合中整合用于并行数据处理的框架,从而使客户易于访问这些服务和部署其程序。但是,当前使用的处理框架是为静态,均匀集群设置而设计的,而无视云的特定性质。因此,分配的计算资源可能不足以在提交的工作的大部分中不足,并不必要地增加处理时间和成本。在本文中,我们讨论了在云中有效并行数据处理的机遇和挑战,并介绍了我们的研究项目Nephele。 Nephele是第一个数据处理框架,可以明确利用当今Iaas云提供的动态资源分配,即任务调度和执行。处理作业的特定任务可以分配给不同类型的虚拟机,这些虚拟机在作业执行过程中会自动实例化和终止。基于这个新框架,我们在Iaas云系统上对MAPReduce启发的处理作业进行了扩展评估,并将结果与​​流行的数据处理框架Hadoop进行了比较。
In recent years ad hoc parallel data processing has emerged to be one of the killer applications for Infrastructure-as-a-Service (IaaS) clouds. Major Cloud computing companies have started to integrate frameworks for parallel data processing in their product portfolio, making it easy for customers to access these services and to deploy their programs. However, the processing frameworks which are currently used have been designed for static, homogeneous cluster setups and disregard the particular nature of a cloud. Consequently, the allocated compute resources may be inadequate for big parts of the submitted job and unnecessarily increase processing time and cost. In this paper, we discuss the opportunities and challenges for efficient parallel data processing in clouds and present our research project Nephele. Nephele is the first data processing framework to explicitly exploit the dynamic resource allocation offered by today's IaaS clouds for both, task scheduling and execution. Particular tasks of a processing job can be assigned to different types of virtual machines which are automatically instantiated and terminated during the job execution. Based on this new framework, we perform extended evaluations of MapReduce-inspired processing jobs on an IaaS cloud system and compare the results to the popular data processing framework Hadoop.