Cost-Aware Streaming Workflow Allocation on Geo-Distributed Data Centers

Cost-Aware Streaming Workflow Allocation on Geo-Distributed Data Centers
复制标题

DOI:
10.1109/tc.2016.2595579
复制
发表时间:
2017-02-01
影响因子:
3.7
通讯作者:
Li, Zhenni
Li, Zhenni
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chen, Wuhui;Paik, Incheon;Li, Zhenni

文献摘要

被引文献

相似文献

云计算中的虚拟机(VM)分配问题近年来得到了广泛的研究,文献中提出了许多算法。其中大部分已成功应用于MapReduce等批处理模型;然而,由于以下缺点,它们都不能很好地应用于流式工作流:1)由于数据流生命周期短,无法捕获流式工作流中任务的特征; 2) 大多数算法都基于这样的假设:虚拟机的价格和数据中心 (DC) 之间的流量是静态且固定的。在本文中,我们提出了一种流式工作流分配算法,该算法考虑了流式工作的特点和地理分布式DC之间的价格多样性,以进一步实现流式大数据处理成本最小化的目标。首先,我们基于流式工作流的任务语义和地理分布式DC的价格多样性构建了扩展流式工作流图(ESWG),并将流式工作流分配问题转化为基于ESWG的混合整数线性规划。其次,我们提出了两种基于任务组合和DC组合的启发式算法来减少计算空间,以满足严格的延迟要求。最后,我们的实验结果表明,总成本和执行时间较低,性能显着提升。
The virtual machine (VM) allocation problem in cloud computing has been widely studied in recent years, and many algorithms have been proposed in the literature. Most of them have been successfully applied to batch processing models such as MapReduce; however, none of them can be applied to streaming workflow well because of the following weaknesses: 1) failure to capture the characteristics of tasks in streaming workflow for the short life cycle of data streams; 2) most algorithms are based on the assumptions that the price of VMs and traffic among data centers (DCs) are static and fixed. In this paper, we propose a streaming workflow allocation algorithm that takes into consideration the characteristics of streaming work and the price diversity among geo-distributed DCs, to further achieve the goal of cost minimization for streaming big data processing. First, we construct an extended streaming workflow graph (ESWG) based on the task semantics of streaming workflow and the price diversity of geo-distributed DCs, and the streaming workflow allocation problem is formulated into mixed integer linear programming based on the ESWG. Second, we propose two heuristic algorithms to reduce the computational space based on task combination and DC combination in order to meet the strict latency requirement. Finally, our experimental results demonstrate significant performance gains with lower total cost and execution time.