Digital Commons @ Michigan Tech Digital Commons @ Michigan Tech Towards real-time reinforcement learning control of a wave Towards real-time reinforcement learning control of a wave energy converter energy converter

Digital Commons @ Michigan Tech Digital Commons @ Michigan Tech Towards real-time reinforcement learning control of a wave Towards real-time reinforcement learning control of a wave energy converter energy converter
复制标题

数字共享@密歇根理工大学 数字共享@密歇根理工大学 迈向波浪的实时强化学习控制 迈向波浪能转换器的实时强化学习控制 能源转换器

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
J. Frazer
J. Frazer
中科院分区:
--
文献类型:
--
作者:
Thomas Fischer;Candy M. Herr;M. Burry;J. Frazer

文献摘要

被引文献

相似文献

波浪能转换器(WECs)的能源成本与化石燃料发电站相比还没有竞争力。为了提高波浪能的可行性,有必要制定有效的控制策略,在温和的海况下最大限度地吸收能量,同时限制高浪的运动。由于其基于模型的性质,最先进的控制方案难以处理模型的不确定性,适应系统动力学随时间的变化,并为大型WECs阵列提供实时集中控制。本文介绍了一种替代解决方案来应对这些挑战,即首次将深度强化学习(DRL)应用于WECs的控制。在线性仿真环境下,采用线性模型预测控制,从多个海况收集的数据初始化DRL代理。该智能体在高波高和高周期时优于模型预测控制,但与WEC的谐振周期接近。通过将计算工作从部署时间转移到训练时间,DRL在部署时的计算成本也大大降低。这为将DRL应用于大型WECs阵列提供了信心,实现了规模经济。此外,无模型强化学习可以自主适应系统动态变化,实现容错控制。
: The levellised cost of energy of wave energy converters (WECs) is not competitive with fossil fuel-powered stations yet. To improve the feasibility of wave energy, it is necessary to develop effective control strategies that maximise energy absorption in mild sea states, whilst limiting motions in high waves. Due to their model-based nature, state-of-the-art control schemes struggle to deal with model uncertainties, adapt to changes in the system dynamics with time, and provide real-time centralised control for large arrays of WECs. Here, an alternative solution is introduced to address these challenges, applying deep reinforcement learning (DRL) to the control of WECs for the first time. A DRL agent is initialised from data collected in multiple sea states under linear model predictive control in a linear simulation environment. The agent outperforms model predictive control for high wave heights and periods, but suffers close to the resonant period of the WEC. The computational cost at deployment time of DRL is also much lower by diverting the computational effort from deployment time to training. This provides confidence in the application of DRL to large arrays of WECs, enabling economies of scale. Additionally, model-free reinforcement learning can autonomously adapt to changes in the system dynamics, enabling fault-tolerant control.