Graph partition–based data and task co‐scheduling of scientific workflow in geo‐distributed datacenters

Graph partition–based data and task co‐scheduling of scientific workflow in geo‐distributed datacenters
复制标题

DOI:
10.1002/cpe.5245
复制
发表时间:
2019-04
期刊:
Concurrency and Computation: Practice and Experience
影响因子:
--
通讯作者:
Jinghui Zhang;Jian Chen;Jun Zhan;Jiahui Jin;Aibo Song
Jinghui Zhang;Jian Chen;Jun Zhan;Jiahui Jin;Aibo Song
中科院分区:
其他
文献类型:
--
作者:
Jinghui Zhang;Jian Chen;Jun Zhan;Jiahui Jin;Aibo Song

文献摘要

相似文献

大多数大规模科学工作流程在多个协作数据中心进行,以访问社区范围的资源,同时遵守每个数据中心的非统一资源限制。然而,在地理分布的数据中心之间移动具有预定位置的初始输入数据集和需要布局决策的中间数据集都会阻碍大规模数据密集型科学工作流的有效执行。因此,科学工作流的数据和任务协同调度处理预先放置的初始输入数据集、中间数据集的放置以及每个数据中心的非均匀计算和存储约束等情况,同时最小化跨数据中心的数据传输。由于该调度问题是NP-Hard问题,本文提出了一种基于多层图粗化和去粗化框架的新方法,并结合一种具有独特的图划分驱动修复和局部改进特性的专用混合遗传算法,来调度地理分布的数据中心中的数据密集型科学工作流,并优化跨数据中心的数据传输量。基于四个真实工作流轨迹的大量仿真表明,我们的算法显著减少了地理分布数据的整体传输,并证明了其有效性。
Most large‐scale scientific workflows take place in multiple collaborative datacenters for access to community‐wide resources, while adhering to each datacenter's non‐uniform resource limits. However, moving both initial input datasets with predetermined locations and intermediate datasets needing placement decisions across geo‐distributed datacenters hinders efficient execution of large‐scale data‐intensive scientific workflows. Thus, scientific workflow's data and task co‐scheduling deal with situations such as pre‐placed initial input datasets, placement of intermediate datasets and each datacenter's non‐uniform computation and storage constraint, while minimizing the cross‐datacenter data transfer. Since this scheduling problem is known to be NP‐hard, here, we propose a novel approach, based on the multilevel graph coarsening and uncoarsening framework, together with a specialized hybrid genetic algorithm having distinctive graph partition driven features of repair and local improvement, for scheduling data‐intensive scientific workflows in geo‐distributed datacenters and optimizing the cross‐datacenter data transfer volume. Extensive simulations, based on four real‐world workflow traces, show that our algorithm significantly reduces the overall geo‐distributed data transfer and demonstrate its effectiveness.