Regularizing Action Policies for Smooth Control with Reinforcement Learning

Regularizing Action Policies for Smooth Control with Reinforcement Learning
复制标题

DOI:
10.1109/icra48506.2021.9561138
复制
发表时间:
2020-12
期刊:
2021 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Siddharth Mysore;B. Mabsout;R. Mancuso;Kate Saenko
Siddharth Mysore;B. Mabsout;R. Mancuso;Kate Saenko
中科院分区:
其他
文献类型:
--
作者:
Siddharth Mysore;B. Mabsout;R. Mancuso;Kate Saenko

文献摘要

相似文献

使用深度强化学习(RL)训练的控制器在实际应用中的一个关键问题是,RL策略学习的动作明显缺乏平滑性。这种趋势通常以控制信号振荡的形式出现,并可能导致控制不良,高功耗和系统过度磨损。我们引入了动作策略平滑条件反射(CAPS),这是一种有效而直观的动作策略正则化,它提供了神经网络控制器学习到的状态到动作映射的平滑性的持续改进,反映在消除控制信号中的高频成分上。在一个真实的系统测试中,在四旋翼无人机上的控制器平滑性的改进导致功耗降低了近80%,同时持续训练飞行控制器。项目网站:http://ai.bu.edu/caps
A critical problem with the practical utility of controllers trained with deep Reinforcement Learning (RL) is the notable lack of smoothness in the actions learned by the RL policies. This trend often presents itself in the form of control signal oscillation and can result in poor control, high power consumption, and undue system wear. We introduce Conditioning for Action Policy Smoothness (CAPS), an effective yet intuitive regularization on action policies, which offers consistent improvement in the smoothness of the learned state-to-action mappings of neural network controllers, reflected in the elimination of high-frequency components in the control signal. Tested on a real system, improvements in controller smoothness on a quadrotor drone resulted in an almost 80% reduction in power consumption while consistently training flight-worthy controllers. Project website: http://ai.bu.edu/caps