A Bayesian Risk Approach to MDPs with Parameter Uncertainty

A Bayesian Risk Approach to MDPs with Parameter Uncertainty
复制标题

具有参数不确定性的 MDP 的贝叶斯风险方法

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Enlu Zhou
Enlu Zhou
中科院分区:
--
文献类型:
--
作者:
Yifan Lin;Yuxuan Ren;Enlu Zhou

文献摘要

被引文献

相似文献

我们考虑马尔可夫决策过程(mdp),其中分布参数,如转移概率,是未知的,并从数据中估计。解决参数不确定性的流行的分布鲁棒方法有时可能过于保守。在本文中,我们提出了一种具有参数不确定性的mdp的贝叶斯风险方法,其中风险函数以嵌套形式应用于每个时间阶段未知参数的贝叶斯后验分布的预期贴现总成本。所提出的方法对参数不确定性的风险态度提供了更大的灵活性,并考虑了未来时间阶段数据的可用性。对于有限视界mdp,我们证明了一种基于上置信度界(UCB)的自适应采样算法可以有效地求解动态规划方程。对于无限视界mdp,我们提出了一个风险调整的Bellman算子,并证明了该算子是一个收缩映射,该映射导致贝叶斯风险公式的最优值函数。在有限视界的情况下,我们在库存控制问题和路径规划问题上证明了我们提出的算法的经验性能。
We consider Markov Decision Processes (MDPs) where distributional parameters, such as transition probabilities, are unknown and estimated from data. The popular distributionally robust approach to addressing the parameter uncertainty can some-times be overly conservative. In this paper, we propose a Bayesian risk approach to MDPs with parameter uncertainty, where a risk functional is applied in nested form to the expected discounted total cost with respect to the Bayesian posterior distributions of the unknown parameters in each time stage. The proposed approach provides more flexibility of risk attitudes towards parameter uncertainty and takes into account the availability of data in future time stages. For the finite-horizon MDPs, we show the dynamic programming equations can be solved efficiently with an upper confidence bound (UCB) based adaptive sampling algorithm. For the infinite-horizon MDPs, we propose a risk-adjusted Bellman operator and show the proposed operator is a contraction mapping that leads to the optimal value function to the Bayesian risk formulation. We demonstrate the empirical performance of our proposed algorithms in the finite-horizon case on an inventory control problem and a path planning problem.