Big Data Aware Virtual Machine Placement in Cloud Data Centers

Big Data Aware Virtual Machine Placement in Cloud Data Centers
复制标题

DOI:
10.1145/3148055.3148057
复制
发表时间:
2017-12
期刊:
Proceedings of the Fourth IEEE/ACM International Conference on Big Data Computing, Applications and Technologies
影响因子:
--
通讯作者:
Logan Hall;B. Harris;Erica Tomes;Nihat Altiparmak
Logan Hall;B. Harris;Erica Tomes;Nihat Altiparmak
中科院分区:
其他
文献类型:
--
作者:
Logan Hall;B. Harris;Erica Tomes;Nihat Altiparmak

文献摘要

相似文献

虽然处理大数据的洞察力继续改变着社会,但收集这些数据的速度越来越快,使得私人集群中的处理变得过时。大量的大数据已经驻留在云中,云基础设施为大数据处理应用程序的计算和I/O需求提供了可扩展的平台。虚拟化被用作云中的基础技术;然而,现有的虚拟机放置技术不考虑基础设施的数据复制和I/O瓶颈,从而产生次优的数据检索时间。本文的目标是在云中高效地处理大数据,并提出了新的虚拟机放置技术,通过考虑数据复制,存储性能和网络带宽来最大限度地减少数据检索时间。首先提出了一种基于整数规划的虚拟机优化布局算法,然后提出了两种低成本的数据和能量感知的虚拟机布局算法。我们提出的算法进行了比较,通过广泛的评估与最佳和现有的算法。实验结果为我们提出的解决方案在性能和能源方面的优越性提供了强有力的证据,并清楚地概述了大数据感知虚拟机放置对于高效处理云中大型数据集的重要性。
While society continues to be transformed by insights from processing big data, the increasing rate at which this data is gathered is making processing in private clusters obsolete. A vast amount of big data already resides in the cloud, and cloud infrastructures provide a scalable platform for both the computational and I/O needs of big data processing applications. Virtualization is used as a base technology in the cloud; however, existing virtual machine placement techniques do not consider data replication and I/O bottlenecks of the infrastructure, yielding sub-optimal data retrieval times. This paper targets efficient big data processing in the cloud and proposes novel virtual machine placement techniques, which minimize data retrieval time by considering data replication, storage performance, and network bandwidth. We first present an integer-programming based optimal virtual machine placement algorithm and then propose two low cost data- and energy-aware virtual machine placement heuristics. Our proposed heuristics are compared with optimal and existing algorithms through extensive evaluation. Experimental results provide strong indications for the superiority of our proposed solutions in both performance and energy, and clearly outline the importance of big data aware virtual machine placement for efficient processing of large datasets in the cloud.