Hybrid control for combining model-based and model-free reinforcement learning

Hybrid control for combining model-based and model-free reinforcement learning
复制标题

DOI:
10.1177/02783649221083331
复制
发表时间:
2022-06
期刊:
The International Journal of Robotics Research
影响因子:
--
通讯作者:
Allison Pinosky;Ian Abraham;Alexander Broad;B. Argall;T. Murphey
Allison Pinosky;Ian Abraham;Alexander Broad;B. Argall;T. Murphey
中科院分区:
其他
文献类型:
--
作者:
Allison Pinosky;Ian Abraham;Alexander Broad;B. Argall;T. Murphey

文献摘要

被引文献

相似文献

我们开发了一种方法,以提高机器人系统的学习能力相结合的学习预测模型与基于经验的状态-动作策略映射。预测模型提供了对任务和动态的理解,而基于经验(无模型)的策略映射编码了覆盖计划行动的有利行动。我们将系统地结合基于模型和无模型学习方法的方法称为混合学习。我们的方法可以有效地学习运动技能,并提高预测模型和基于经验的策略的性能。此外,我们的方法可以使用任何非策略强化学习方法更新策略(基于模型和无模型)。我们得出一个确定性的方法,混合学习的学习模式之间的最佳切换。我们调整我们的方法,放松一些在原来的推导中的关键假设的随机变化。我们的确定性和随机性的变化进行了测试,各种机器人控制基准任务的模拟,以及硬件操作任务。我们扩展我们的方法与模仿学习方法,通过演示提供经验,我们测试扩展的能力与现实世界的拾取和放置任务。结果表明,我们的方法是能够提高性能和样本效率的学习运动技能在各种实验领域。
We develop an approach to improve the learning capabilities of robotic systems by combining learned predictive models with experience-based state-action policy mappings. Predictive models provide an understanding of the task and the dynamics, while experience-based (model-free) policy mappings encode favorable actions that override planned actions. We refer to our approach of systematically combining model-based and model-free learning methods as hybrid learning. Our approach efficiently learns motor skills and improves the performance of predictive models and experience-based policies. Moreover, our approach enables policies (both model-based and model-free) to be updated using any off-policy reinforcement learning method. We derive a deterministic method of hybrid learning by optimally switching between learning modalities. We adapt our method to a stochastic variation that relaxes some of the key assumptions in the original derivation. Our deterministic and stochastic variations are tested on a variety of robot control benchmark tasks in simulation as well as a hardware manipulation task. We extend our approach for use with imitation learning methods, where experience is provided through demonstrations, and we test the expanded capability with a real-world pick-and-place task. The results show that our method is capable of improving the performance and sample efficiency of learning motor skills in a variety of experimental domains.