Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems

Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
复制标题

DOI:
10.48550/arxiv.2402.01147
复制
发表时间:
2024-02
期刊:
--
影响因子:
--
通讯作者:
Neharika Jali;Guannan Qu;Weina Wang;Gauri Joshi
Neharika Jali;Guannan Qu;Weina Wang;Gauri Joshi
中科院分区:
其他
文献类型:
--
作者:
Neharika Jali;Guannan Qu;Weina Wang;Gauri Joshi

文献摘要

相似文献

我们考虑有效地将到达中央队列的作业路由到异构服务器系统的问题。与同构系统不同,阈值策略(当队列长度超过某个阈值时将作业路由到较慢的服务器)对于一快一慢双服务器系统来说是最优的。但是多服务器系统的最优策略是未知的,很难找到。虽然强化学习(RL)已经被认为在这种情况下具有学习策略的巨大潜力,但我们的问题具有指数级大的状态空间大小,使得标准RL效率低下。在这项工作中,我们提出了ACHQ,这是一种高效的基于策略梯度的算法,具有低维软阈值策略参数化,利用底层排队结构。我们提供了一般情况下的平稳点收敛保证,并在低维参数化的情况下证明了对于两个服务器的特殊情况ACHQ收敛到近似全局最优。仿真表明,与路由到最快可用服务器的贪婪策略相比,预期响应时间最多可提高30%。
We consider the problem of efficiently routing jobs that arrive into a central queue to a system of heterogeneous servers. Unlike homogeneous systems, a threshold policy, that routes jobs to the slow server(s) when the queue length exceeds a certain threshold, is known to be optimal for the one-fast-one-slow two-server system. But an optimal policy for the multi-server system is unknown and non-trivial to find. While Reinforcement Learning (RL) has been recognized to have great potential for learning policies in such cases, our problem has an exponentially large state space size, rendering standard RL inefficient. In this work, we propose ACHQ, an efficient policy gradient based algorithm with a low dimensional soft threshold policy parameterization that leverages the underlying queueing structure. We provide stationary-point convergence guarantees for the general case and despite the low-dimensional parameterization prove that ACHQ converges to an approximate global optimum for the special case of two servers. Simulations demonstrate an improvement in expected response time of up to ~30% over the greedy policy that routes to the fastest available server.