Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs

Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Marek Petrik;R. Russel
Marek Petrik;R. Russel
中科院分区:
其他
文献类型:
--
作者:
Marek Petrik;R. Russel

文献摘要

相似文献

鲁棒MDP(RMDP)可以用于计算强化学习中具有可证明的最坏情况保证的策略。RMDP解决方案的质量和鲁棒性由模糊集-合理转移概率集-决定,模糊集通常构造为多维置信区域。现有的方法构造模糊集的置信区域使用浓度不等式,导致过于保守的解决方案。本文提出了一种新的范例,可以实现更好的解决方案,具有相同的鲁棒性保证,而不使用置信区域作为模糊集。为了结合先验知识,我们的算法使用贝叶斯推理优化歧义集的大小和位置。我们的理论分析表明,所提出的方法的安全性,和实证结果表明其实际的承诺。
Robust MDPs (RMDPs) can be used to compute policies with provable worst-case guarantees in reinforcement learning. The quality and robustness of an RMDP solution are determined by the ambiguity set---the set of plausible transition probabilities---which is usually constructed as a multi-dimensional confidence region. Existing methods construct ambiguity sets as confidence regions using concentration inequalities which leads to overly conservative solutions. This paper proposes a new paradigm that can achieve better solutions with the same robustness guarantees without using confidence regions as ambiguity sets. To incorporate prior knowledge, our algorithms optimize the size and position of ambiguity sets using Bayesian inference. Our theoretical analysis shows the safety of the proposed method, and the empirical results demonstrate its practical promise.