MARS: Malleable Actor-Critic Reinforcement Learning Scheduler
MARS: Malleable Actor-Critic Reinforcement Learning Scheduler
复制标题
DOI:
10.1109/ipccc55026.2022.9894315
复制
发表时间:
2020-05
期刊:
影响因子:
--
通讯作者:
Betis Baheri;Jake Tronge;B. Fang;Ang Li;V. Chaudhary;Qiang Guan
中科院分区:
文献类型:
--
作者:
Betis Baheri;Jake Tronge;B. Fang;Ang Li;V. Chaudhary;Qiang Guan
In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARSensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the predefined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARSupdates the Deep Neural Network (DNN) model based on the reward. MARSis designed to optimize the existing models through reinforcement mechanisms. MARSadapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self-learning deep neural network model at run-time. We evaluate MARSwith different real-world workflow traces. MARS can achieve 5%–60% increased performance compared to the state-of-the-art approaches.