Deterministic Sequencing of Exploration and Exploitation for Reinforcement Learning

Deterministic Sequencing of Exploration and Exploitation for Reinforcement Learning
复制标题

DOI:
10.1109/cdc51059.2022.9992857
复制
发表时间:
2022-09
期刊:
2022 IEEE 61st Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
P. Gupta;Vaibhav Srivastava
P. Gupta;Vaibhav Srivastava
中科院分区:
其他
文献类型:
--
作者:
P. Gupta;Vaibhav Srivastava

文献摘要

被引文献

相似文献

我们提出了基于模型的RL问题的确定性探索和开发排序(DSEE)算法,该算法具有交错的探索和开发时期,旨在同时学习系统模型,即,马尔可夫决策过程(MDP)和相关的最优策略。在探索期间,DSEE探索环境并更新预期奖励和转移概率的估计。在利用过程中,使用期望回报和转移概率的最新估计来获得高概率的鲁棒策略。我们设计了探索和利用时期的长度,使得累积的后悔随着时间的次线性函数而增长。
We propose Deterministic Sequencing of Exploration and Exploitation (DSEE) algorithm with interleaving exploration and exploitation epochs for model-based RL problems that aim to simultaneously learn the system model, i.e., a Markov decision process (MDP), and the associated optimal policy. During exploration, DSEE explores the environment and updates the estimates for expected reward and transition probabilities. During exploitation, the latest estimates of the expected reward and transition probabilities are used to obtain a robust policy with high probability. We design the lengths of the exploration and exploitation epochs such that the cumulative regret grows as a sub-linear function of time.