Robust and Adaptive Planning under Model Uncertainty

Robust and Adaptive Planning under Model Uncertainty
复制标题

DOI:
10.1609/icaps.v29i1.3505
复制
发表时间:
2019-01
期刊:
--
影响因子:
--
通讯作者:
Apoorva Sharma;James Harrison;Matthew W. Tsao;M. Pavone
Apoorva Sharma;James Harrison;Matthew W. Tsao;M. Pavone
中科院分区:
其他
文献类型:
--
作者:
Apoorva Sharma;James Harrison;Matthew W. Tsao;M. Pavone

文献摘要

被引文献

相似文献

模型不确定性下的规划是许多决策和学习应用中的一个基本问题。在本文中,我们提出了鲁棒自适应蒙特卡罗规划(RAMCP)算法,它允许计算风险敏感的贝叶斯自适应政策,最佳权衡探索,剥削和鲁棒性。RAMCP制定的风险敏感规划问题作为一个两个玩家的零和游戏,其中一个对手扰动代理的信念模型。我们介绍两个版本的RAMCP算法。第一,RAMCP-F,收敛到一个最佳的risksensitive政策,而不必重建搜索树的基础上的信念模型被扰动。第二个版本,RAMCP-I,提高了计算效率的成本失去理论保证,但产生的经验结果相比,RAMCP-F。RAMCP是证明了一个N拉多臂土匪问题,以及病人的治疗方案。
Planning under model uncertainty is a fundamental problem across many applications of decision making and learning. In this paper, we propose the Robust Adaptive Monte Carlo Planning (RAMCP) algorithm, which allows computation of risk-sensitive Bayes-adaptive policies that optimally trade off exploration, exploitation, and robustness. RAMCP formulates the risk-sensitive planning problem as a two-player zero-sum game, in which an adversary perturbs the agent’s belief over the models. We introduce two versions of the RAMCP algorithm. The first, RAMCP-F, converges to an optimal risksensitive policy without having to rebuild the search tree as the underlying belief over models is perturbed. The second version, RAMCP-I, improves computational efficiency at the cost of losing theoretical guarantees, but is shown to yield empirical results comparable to RAMCP-F. RAMCP is demonstrated on an n-pull multi-armed bandit problem, as well as a patient treatment scenario.