MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler
复制标题

DOI:
10.1109/ipccc55026.2022.9894315
复制
发表时间:
2020-05
期刊:
2022 IEEE International Performance, Computing, and Communications Conference (IPCCC)
影响因子:
--
通讯作者:
Betis Baheri;Jake Tronge;B. Fang;Ang Li;V. Chaudhary;Qiang Guan
Betis Baheri;Jake Tronge;B. Fang;Ang Li;V. Chaudhary;Qiang Guan
中科院分区:
其他
文献类型:
--
作者:
Betis Baheri;Jake Tronge;B. Fang;Ang Li;V. Chaudhary;Qiang Guan

文献摘要

相似文献

在本文中,我们介绍了 MARS,这是一种基于成本感知、灵活的强化学习方法的 HPC 云基础设施的新型调度系统,可作为下一代 HPC 云资源管理器的中间层。 MARSense 将来自启发式工作负载的预训练模型进行组合,并决定最具成本效益的优化策略。整个工作流应用程序将被拆分为多个可优化的相互依赖的子任务,然后根据预定义的资源管理计划,在执行计划任务后生成奖励。最后,MARS根据奖励更新深度神经网络(DNN)模型。 MAR旨在通过强化机制优化现有模型。 MARS适应工作流应用的动态性,在运行时预先构建的调度策略(回填、SJF等)和自学习深度神经网络模型中选择最具成本效益的调度解决方案。我们使用不同的现实世界工作流程轨迹来评估 MARS。与最先进的方法相比,MARS 的性能可以提高 5%–60%。
In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARSensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the predefined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARSupdates the Deep Neural Network (DNN) model based on the reward. MARSis designed to optimize the existing models through reinforcement mechanisms. MARSadapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self-learning deep neural network model at run-time. We evaluate MARSwith different real-world workflow traces. MARS can achieve 5%–60% increased performance compared to the state-of-the-art approaches.