Indexability is Not Enough for Whittle: Improved, Near-Optimal Algorithms for Restless Bandits

Indexability is Not Enough for Whittle: Improved, Near-Optimal Algorithms for Restless Bandits
复制标题

对于 Whittle 来说,可索引性还不够:针对不安分强盗的改进的、近乎最优的算法

DOI:
--
复制
发表时间:
2022
期刊:
Adaptive Agents and Multi-Agent Systems
影响因子:
--
通讯作者:
Milind Tambe
Milind Tambe
中科院分区:
--
文献类型:
--
作者:
Abheek Ghosh;Dheeraj M. Nagaraj;Manish Jain;Milind Tambe

文献摘要

参考文献

被引文献

相似文献

我们研究了通过多种行动来规划不安分的多臂强盗 (RMAB) 的问题。这是多代理系统的流行模型,具有多通道通信、监控和机器维护任务以及医疗保健等应用。基于拉格朗日松弛的 Whittle 指数策略因其简单性和在某些条件下接近最优而被广泛应用于这些环境中。在这项工作中,我们首先表明,即使 RMAB 可索引,Whittle 索引策略也可能在简单且实际相关的 RMAB 设置中失败。我们讨论为什么最优性保证失败以及为什么渐近最优性可能不能很好地转化为实际相关的规划范围。然后,我们提出了一种基于平均场方法的替代规划算法,该算法可以证明且有效地获得具有大量臂的接近最优策略,而无需 Whittle 指数策略所需的严格结构假设。这借鉴了现有研究的想法,并进行了一些改进:我们的方法是无超参数的,并且我们提供了改进的非渐近分析,其具有:(a)不需要外生超参数和对已知问题参数更严格的多项式依赖; (b) 高概率界限,表明该政策的奖励是可靠的; (c) 匹配该算法相对于臂数的次优下限,从而证明我们的边界的紧密性。我们广泛的实验分析表明,平均场方法与其他基线相匹配或优于其他基线。
We study the problem of planning restless multi-armed bandits (RMABs) with multiple actions. This is a popular model for multi-agent systems with applications like multi-channel communication, monitoring and machine maintenance tasks, and healthcare. Whittle index policies, which are based on Lagrangian relaxations, are widely used in these settings due to their simplicity and near-optimality under certain conditions. In this work, we first show that Whittle index policies can fail in simple and practically relevant RMAB settings, even when the RMABs are indexable. We discuss why the optimality guarantees fail and why asymptotic optimality may not translate well to practically relevant planning horizons. We then propose an alternate planning algorithm based on the mean-field method, which can provably and efficiently obtain near-optimal policies with a large number of arms, without the stringent structural assumptions required by the Whittle index policies. This borrows ideas from existing research with some improvements: our approach is hyper-parameter free, and we provide an improved non-asymptotic analysis which has: (a) no requirement for exogenous hyper-parameters and tighter polynomial dependence on known problem parameters; (b) high probability bounds which show that the reward of the policy is reliable; and (c) matching sub-optimality lower bounds for this algorithm with respect to the number of arms, thus demonstrating the tightness of our bounds. Our extensive experimental analysis shows that the mean-field approach matches or outperforms other baselines.
排队控制和资产管理的可索引性的一般概念
DOI: 10.1214/10-aap705
发表时间: 2011
期刊: The Annals of Applied Probability
影响因子: --
作者:
Glazebrook K
通讯作者: Glazebrook K
灵活电力需求的Kullback-Leibler-二次最优控制
DOI: 10.1109/cdc40024.2019.9029512
发表时间: 2019
期刊: IEEE Conference Decision and Control
影响因子: --
作者:
Cammardella, Neil;Busic, Ana;Ji, Yuting;Meyn, Sean
通讯作者: Meyn, Sean
DOI: 10.1137/18m1196479
发表时间: 2018-02
期刊: SIAM J. Control. Optim.
影响因子: --
作者:
Beatrice Acciaio;Julio D. Backhoff Veraguas;R. Carmona
通讯作者: Beatrice Acciaio;Julio D. Backhoff Veraguas;R. Carmona
DOI: 10.1136/bmjopen-2019-030473
发表时间: 2019-05-01
期刊: BMJ OPEN
影响因子: 2.9
作者:
Erguera, Xavier A.;Johnson, Mallory O.;Saberi, Parya
通讯作者: Saberi, Parya