Physics-informed Dyna-style model-based deep reinforcement learning for dynamic control

Physics-informed Dyna-style model-based deep reinforcement learning for dynamic control
复制标题

DOI:
10.1098/rspa.2021.0618
复制
发表时间:
2021-07
期刊:
Proceedings of the Royal Society A
影响因子:
--
通讯作者:
Xin-Yang Liu;Jian-Xun Wang
Xin-Yang Liu;Jian-Xun Wang
中科院分区:
其他
文献类型:
--
作者:
Xin-Yang Liu;Jian-Xun Wang

文献摘要

被引文献

相似文献

通过学习环境的预测模型,基于模型的强化学习(MBRL)被认为比无模型算法具有更高的样本效率。然而,MBRL 的性能高度依赖于学习模型的质量,该模型通常以黑盒方式构建,在数据分布之外的预测准确性可能较差。学习模型的缺陷可能会阻碍策略的充分优化。尽管已经提出了一些基于不确定性分析的补救措施来缓解这个问题,但模型偏差仍然对 MBRL 提出了巨大的挑战。在这项工作中,我们建议利用环境基础物理的先验知识,其中控制定律(部分)已知。特别是,我们开发了一个基于物理的 MBRL 框架,其中控制方程和物理约束用于为模型学习和策略搜索提供信息。通过结合环境的先验信息,可以显着提高学习模型的质量,同时显着减少所需的与环境的交互,从而获得更好的样本效率和学习性能。其有效性和优点已在一些经典控制问题中得到证明,其中环境由规范常微分方程/偏微分方程控制。
Model-based reinforcement learning (MBRL) is believed to have much higher sample efficiency compared with model-free algorithms by learning a predictive model of the environment. However, the performance of MBRL highly relies on the quality of the learned model, which is usually built in a black-box manner and may have poor predictive accuracy outside of the data distribution. The deficiencies of the learned model may prevent the policy from being fully optimized. Although some uncertainty analysis-based remedies have been proposed to alleviate this issue, model bias still poses a great challenge for MBRL. In this work, we propose to leverage the prior knowledge of underlying physics of the environment, where the governing laws are (partially) known. In particular, we developed a physics-informed MBRL framework, where governing equations and physical constraints are used to inform the model learning and policy search. By incorporating the prior information of the environment, the quality of the learned model can be notably improved, while the required interactions with the environment are significantly reduced, leading to better sample efficiency and learning performance. The effectiveness and merit have been demonstrated over a handful of classic control problems, where the environments are governed by canonical ordinary/partial differential equations.