RLScheduler: An Automated HPC Batch Job Scheduler Using Reinforcement Learning

RLScheduler: An Automated HPC Batch Job Scheduler Using Reinforcement Learning
复制标题

DOI:
10.1109/sc41405.2020.00035
复制
发表时间:
2020-11
期刊:
SC20: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Di Zhang;Dong Dai;Youbiao He;F. Bao;Bing Xie
Di Zhang;Dong Dai;Youbiao He;F. Bao;Bing Xie
中科院分区:
其他
文献类型:
--
作者:
Di Zhang;Dong Dai;Youbiao He;F. Bao;Bing Xie

文献摘要

被引文献

相似文献

当今的高性能计算(HPC)平台仍然以批处理作业为主。因此,有效的批作业调度是获得高系统效率的关键。现有的HPC批处理作业调度器通常利用启发式优先级函数来对作业进行优先级排序和调度。但是,一旦由专家配置和部署,这样的优先级函数很难适应作业负载、优化目标或系统设置的变化,当发生变化时,可能导致系统效率降低。为了解决这个根本问题,我们提出了RLScheduler,这是一个基于强化学习的自动化HPC批处理作业调度器。RLSTOM依赖于最少的人工干预或专家知识,但可以通过自己的连续“试错”学习高质量的调度策略。我们在RLSTOM中引入了一种新的基于核的神经网络结构和轨迹过滤机制,以改善和稳定学习过程。通过广泛的评估,我们确认RLSTOM可以学习高质量的调度策略,对各种工作负载和各种优化目标,相对较低的计算成本。此外,我们表明,即使应用于看不见的工作负载时,学习的模型也能稳定地执行,使其适用于生产使用。
Today’s high-performance computing (HPC) platforms are still dominated by batch jobs. Accordingly, effective batch job scheduling is crucial to obtain high system efficiency. Existing HPC batch job schedulers typically leverage heuristic priority functions to prioritize and schedule jobs. But, once configured and deployed by the experts, such priority functions can hardly adapt to the changes of job loads, optimization goals, or system settings, potentially leading to degraded system efficiency when changes occur. To address this fundamental issue, we present RLScheduler, an automated HPC batch job scheduler built on reinforcement learning. RLScheduler relies on minimal manual interventions or expert knowledge, but can learn high-quality scheduling policies via its own continuous ‘trial and error’. We introduce a new kernel-based neural network structure and trajectory filtering mechanism in RLScheduler to improve and stabilize the learning process. Through extensive evaluations, we confirm that RLScheduler can learn high-quality scheduling policies towards various workloads and various optimization goals with relatively low computation cost. Moreover, we show that the learned models perform stably even when applied to unseen workloads, making them practical for production use.