Active Flow Control for Bluff Body Drag Reduction Using Reinforcement Learning with Partial Measurements

Active Flow Control for Bluff Body Drag Reduction Using Reinforcement Learning with Partial Measurements
复制标题

DOI:
10.1017/jfm.2024.69
复制
发表时间:
2023-07
影响因子:
3.7
通讯作者:
C. Xia;Junjie Zhang;E. Kerrigan;Georgios Rigas
C. Xia;Junjie Zhang;E. Kerrigan;Georgios Rigas
中科院分区:
工程技术2区
文献类型:
--
作者:
C. Xia;Junjie Zhang;E. Kerrigan;Georgios Rigas

文献摘要

相似文献

摘要:通过强化学习 (RL) 进行减阻的主动流动控制是在具有涡流脱落的层流状态下的二维方形钝体之后进行的。由神经网络参数化的控制器经过训练,可以驱动两个吹气和吸力喷嘴来操纵不稳定的流动。具有完全可观测性(尾流中的传感器)的强化学习成功地发现了一种控制策略,通过抑制尾流中的涡流脱落来减少阻力。然而,当控制器接受部分测量(身体上的传感器)训练时,会观察到不可忽略的性能下降(阻力减少$\sim$50%)。为了减轻这种影响,我们提出了一种节能、动态、最大熵 RL 控制方案。首先,提出了一种基于能源效率的奖励函数,以优化控制器的能耗,同时最大限度地减少阻力。其次,控制器使用由当前和过去的测量和动作组成的增强状态进行训练,可以将其表示为非线性自回归外生模型,以缓解部分可观测性问题。第三,使用最大熵强化学习算法(软演员批评家和截断分位数批评家),以样本有效的方式促进探索和利用,并在具有挑战性的部分测量情况下发现接近最优的策略。仅使用身体后部的表面压力测量即可在近尾流中实现涡流脱落的稳定,从而实现与尾流传感器的情况类似的阻力减少。所提出的方法为使用部分测量来实现实际配置的动态流量控制开辟了新途径。
Abstract Active flow control for drag reduction with reinforcement learning (RL) is performed in the wake of a two-dimensional square bluff body at laminar regimes with vortex shedding. Controllers parametrised by neural networks are trained to drive two blowing and suction jets that manipulate the unsteady flow. The RL with full observability (sensors in the wake) discovers successfully a control policy that reduces the drag by suppressing the vortex shedding in the wake. However, a non-negligible performance degradation ($\sim$50 % less drag reduction) is observed when the controller is trained with partial measurements (sensors on the body). To mitigate this effect, we propose an energy-efficient, dynamic, maximum entropy RL control scheme. First, an energy-efficiency-based reward function is proposed to optimise the energy consumption of the controller while maximising drag reduction. Second, the controller is trained with an augmented state consisting of both current and past measurements and actions, which can be formulated as a nonlinear autoregressive exogenous model, to alleviate the partial observability problem. Third, maximum entropy RL algorithms (soft actor critic and truncated quantile critics) that promote exploration and exploitation in a sample-efficient way are used, and discover near-optimal policies in the challenging case of partial measurements. Stabilisation of the vortex shedding is achieved in the near wake using only surface pressure measurements on the rear of the body, resulting in drag reduction similar to that in the case with wake sensors. The proposed approach opens new avenues for dynamic flow control using partial measurements for realistic configurations.