On Optimality of Myopic Policy for Restless Multi-Armed Bandit Problem: An Axiomatic Approach

On Optimality of Myopic Policy for Restless Multi-Armed Bandit Problem: An Axiomatic Approach
复制标题

DOI:
10.1109/tsp.2011.2170684
复制
发表时间:
2012
影响因子:
5.4
通讯作者:
Kehao Wang;Lin Chen
Kehao Wang;Lin Chen
中科院分区:
工程技术1区
文献类型:
--
作者:
Kehao Wang;Lin Chen

文献摘要

被引文献

相似文献

由于其在许多工程问题中的应用,不安分的多臂强盗(RMAB)问题是随机决策理论中的一个重要问题。然而,解决RMAB问题是众所周知的PSPACE困难的,由于指数计算复杂度,最优策略通常难以处理。另一种自然的办法是寻求简单而又容易执行的短视政策。本文对RMAB问题的近视策略的最优性进行了一般性研究。更具体地说,我们开发了三个公理表征一个家庭的通用和实际上重要的功能称为定期功能。通过数学分析的基础上开发的公理,我们建立了封闭形式的条件下,近视的政策是最优的保证。公理分析还阐明了短视政策的重要工程含义,包括勘探和开采之间的内在权衡。最后通过一个案例说明了所得结果在分析一类多信道机会接入的RMAB问题中的应用。
Due to its application in numerous engineering problems, the restless multi-armed bandit (RMAB) problem is of fundamental importance in stochastic decision theory. However, solving the RMAB problem is well known to be PSPACE-hard, with the optimal policy usually intractable due to the exponential computation complexity. A natural alternative approach is to seek simple myopic policies which are easy to implement. This paper presents a generic study on the optimality of the myopic policy for the RMAB problem. More specifically, we develop three axioms characterizing a family of generic and practically important functions termed as regular functions. By performing a mathematical analysis based on the developed axioms, we establish the closed-form conditions under which the myopic policy is guaranteed to be optimal. The axiomatic analysis also illuminates important engineering implications of the myopic policy including the intrinsic tradeoff between exploration and exploitation. A case study is then presented to illustrate the application of the derived results in analyzing a class of RMAB problems arising from multi-channel opportunistic access.