Asymmetric DQN for partially observable reinforcement learning

Asymmetric DQN for partially observable reinforcement learning
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Andrea Baisero;Brett Daley;Chris Amato
Andrea Baisero;Brett Daley;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Andrea Baisero;Brett Daley;Chris Amato

文献摘要

相似文献

模拟部分可观察环境中的离线训练允许强化学习方法通​​过一种称为不对称的机制来利用特权状态信息。如果使用得当,这些特权信息有可能大大提高最佳收敛特性。然而,当前非对称强化学习的研究本质上往往是启发式的,与基础理论或理论保证很少有联系,并且主要通过经验评估进行测试。在这项工作中,我们开发了非对称策略迭代理论(一种基于精确模型的动态规划求解方法),然后应用松弛,最终产生非对称 DQN(一种无模型深度强化学习算法)。我们的理论发现得到了在具有大量部分可观察性并且需要信息收集策略和记忆的环境中进行的实证实验的补充和验证。
Offline training in simulated partially observable environments allows reinforcement learning methods to exploit privileged state information through a mechanism known as asymmetry. Such privileged information has the potential to greatly improve the optimal convergence properties, if used appropriately. However, current research in asymmetric reinforcement learning is often heuristic in nature, with few connections to underlying theory or theoretical guarantees, and is primarily tested through empirical evaluation. In this work, we develop the theory of Asymmetric Policy Iteration , an exact model-based dynamic programming solution method, and then apply relaxations which eventually result in Asymmetric DQN , a model-free deep reinforcement learning algorithm. Our theoretical findings are complemented and validated by empirical experimentation performed in environments which exhibit significant amounts of partial observability, and require both information gathering strategies and memorization.