Interpretable policies for reinforcement learning by genetic programming

Interpretable policies for reinforcement learning by genetic programming
复制标题

DOI:
10.1016/j.engappai.2018.09.007
复制
发表时间:
2018-11-01
影响因子:
8
通讯作者:
Runkler, Thomas A.
Runkler, Thomas A.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hein, Daniel;Udluft, Steffen;Runkler, Thomas A.

文献摘要

被引文献

相似文献

寻找可解释的强化学习策略具有很高的学术和工业兴趣。特别是对于工业系统,领域专家更有可能部署自主学习的控制器,如果他们是可以理解的,方便评估。基本代数方程应该满足这些要求,只要它们被限制在足够的复杂性。在这里,我们介绍了遗传编程强化学习(GPRL)方法的基础上,基于模型的批量强化学习和遗传编程,自主学习政策方程从预先存在的默认状态动作轨迹样本。GPRL相比,一个简单的方法,利用遗传编程的符号回归,产生的政策模仿现有的表现良好,但不可解释的政策。三个强化学习基准上的实验,即,山地车,推车杆平衡,和工业基准,证明了我们的GPRL方法相比,符号回归方法的优越性。GPRL能够从预先存在的默认轨迹数据中产生性能良好的可解释的强化学习策略。
The search for interpretable reinforcement learning policies is of high academic and industrial interest. Especially for industrial systems, domain experts are more likely to deploy autonomously learned controllers if they are understandable and convenient to evaluate. Basic algebraic equations are supposed to meet these requirements, as long as they are restricted to an adequate complexity. Here we introduce the genetic programming for reinforcement learning (GPRL) approach based on model-based batch reinforcement learning and genetic programming, which autonomously learns policy equations from pre-existing default state action trajectory samples. GPRL is compared to a straightforward method which utilizes genetic programming for symbolic regression, yielding policies imitating an existing well-performing, but non-interpretable policy. Experiments on three reinforcement learning benchmarks, i.e., mountain car, cart pole balancing, and industrial benchmark, demonstrate the superiority of our GPRL approach compared to the symbolic regression method. GPRL is capable of producing well-performing interpretable reinforcement learning policies from pre-existing default trajectory data.