NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RL

NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RL
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Khaled Nakhleh;Santosh Ganji;Ping-Chun Hsieh;I.-Hong Hou;S. Shakkottai
Khaled Nakhleh;Santosh Ganji;Ping-Chun Hsieh;I.-Hong Hou;S. Shakkottai
中科院分区:
其他
文献类型:
--
作者:
Khaled Nakhleh;Santosh Ganji;Ping-Chun Hsieh;I.-Hong Hou;S. Shakkottai

文献摘要

被引文献

相似文献

Whittle指数策略是一个强有力的工具,以获得渐近最优解的臭名昭著的棘手问题的不安土匪。然而,找到惠特尔指数仍然是一个困难的问题,许多实际的不安分的强盗与复杂的过渡内核。本文提出了NeurWIN,一个神经Whittle指数网络,旨在通过利用Whittle指数的数学特性来学习任何不安分的土匪的Whittle指数。我们表明,神经网络产生的惠特尔指数也是一个产生最优控制的一组马尔可夫决策问题。这一特性促使我们使用深度强化学习来训练NeurWIN。我们证明了效用的NeurWIN最近研究的三个不安分的土匪问题,通过评估其性能。实验结果表明,NeurWIN的性能明显优于其他RL算法。
Whittle index policy is a powerful tool to obtain asymptotically optimal solutions for the notoriously intractable problem of restless bandits. However, finding the Whittle indices remains a difficult problem for many practical restless bandits with convoluted transition kernels. This paper proposes NeurWIN, a neural Whittle index network that seeks to learn the Whittle indices for any restless bandits by leveraging mathematical properties of the Whittle indices. We show that a neural network that produces the Whittle index is also one that produces the optimal control for a set of Markov decision problems. This property motivates using deep reinforcement learning for the training of NeurWIN. We demonstrate the utility of NeurWIN by evaluating its performance for three recently studied restless bandit problems. Our experiment results show that the performance of NeurWIN is significantly better than other RL algorithms.