Autonomous Learning in a Pseudo-Episodic Physical Environment

Autonomous Learning in a Pseudo-Episodic Physical Environment
复制标题

DOI:
10.1007/s10846-022-01577-5
复制
发表时间:
2022-02
影响因子:
3.3
通讯作者:
Kevin P. T. Haughn;D. Inman
Kevin P. T. Haughn;D. Inman
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kevin P. T. Haughn;D. Inman

文献摘要

相似文献

出于实际考虑,强化学习在应用于物理实验时已被证明是模拟之外的一项艰巨任务。在这里,我们通过仔细的实验​​设计和算法决策,得出了一种可选的无模型强化学习方法,完全在线实现。我们设计了一种强化学习方案,以针对不稳定的一维机械环境实施传统的情景算法。培训计划是完全自主的,整个学习过程中不需要有人在场。我们表明,伪情景技术允许通过非策略演员批评家和经验重播方法进行额外的学习更新。我们表明,在传统训练周期之间加入这些额外的更新可以提高学习的速度和一致性。此外,我们在实验硬件中验证了该过程。在物理环境中,几种算法变体学习得很快,每种算法都超过了基线最大奖励。本研究中的算法是无模型的,仅使用机载传感器在训练期间获得的信息。
Forpractical considerations reinforcement learning has proven to be a difficult task outside of simulation when applied to a physical experiment. Here we derive an optional approach to model free reinforcement learning, achieved entirely online, through careful experimental design and algorithmic decision making. We design a reinforcement learning scheme to implement traditionally episodic algorithms for an unstable 1-dimensional mechanical environment. The training scheme is completely autonomous, requiring no human to be present throughout the learning process. We show that the pseudo-episodic technique allows for additional learning updates with off-policy actor-critic and experience replay methods. We show that including these additional updates between periods of traditional training episodes can improve speed and consistency of learning. Furthermore, we validate the procedure in experimental hardware. In the physical environment, several algorithm variants learned rapidly, each surpassing baseline maximum reward. The algorithms in this research are model free and use only information obtained by an onboard sensor during training.