Active Model Estimation in Markov Decision Processes

Active Model Estimation in Markov Decision Processes
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
--
影响因子:
--
通讯作者:
Jean Tarbouriech;S. Shekhar;Matteo Pirotta;M. Ghavamzadeh;A. Lazaric
Jean Tarbouriech;S. Shekhar;Matteo Pirotta;M. Ghavamzadeh;A. Lazaric
中科院分区:
其他
文献类型:
--
作者:
Jean Tarbouriech;S. Shekhar;Matteo Pirotta;M. Ghavamzadeh;A. Lazaric

文献摘要

被引文献

相似文献

我们研究有效探索的问题是为了学习环境的准确模型,该模型被建模为马尔可夫决策过程(MDP)。对这个问题的有效探索要求代理识别估计模型更困难的区域,然后利用这一知识在那里收集更多样本。本文将这一问题形式化,介绍了学习动态的精确估计的第一个算法,并给出了其样本复杂性分析。虽然这种算法在大样本情况下有很强的保证,但在探索的早期阶段,它的性能往往很差。为了解决这个问题,我们提出了一种基于最大加权熵的算法,这是一种源于常识和理论分析的启发式算法。这里的主要思想是用与转换中的噪声成比例的权重来覆盖整个状态-动作空间。使用一些具有异质噪声的简单区域,我们证明了我们的启发式算法在小样本区域内的性能优于我们的原始算法和最大熵算法,而获得了与原始算法相似的渐近性能。
We study the problem of efficient exploration in order to learn an accurate model of an environment, modeled as a Markov decision process (MDP). Efficient exploration in this problem requires the agent to identify the regions in which estimating the model is more difficult and then exploit this knowledge to collect more samples there. In this paper, we formalize this problem, introduce the first algorithm to learn an $\epsilon$-accurate estimate of the dynamics, and provide its sample complexity analysis. While this algorithm enjoys strong guarantees in the large-sample regime, it tends to have a poor performance in early stages of exploration. To address this issue, we propose an algorithm that is based on maximum weighted entropy, a heuristic that stems from common sense and our theoretical analysis. The main idea here is to cover the entire state-action space with the weight proportional to the noise in the transitions. Using a number of simple domains with heterogeneous noise in their transitions, we show that our heuristic-based algorithm outperforms both our original algorithm and the maximum entropy algorithm in the small sample regime, while achieving similar asymptotic performance as that of the original algorithm.