Learning to school in dense configurations with multi-agent deep reinforcement learning

Learning to school in dense configurations with multi-agent deep reinforcement learning
复制标题

DOI:
10.1088/1748-3190/ac9fb5
复制
发表时间:
2022-11
影响因子:
3.4
通讯作者:
Yi Zhu;Jinhui Pang;Tong Gao;F. Tian
Yi Zhu;Jinhui Pang;Tong Gao;F. Tian
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yi Zhu;Jinhui Pang;Tong Gao;F. Tian

文献摘要

被引文献

相似文献

人们观察到鱼以不同的形态聚集。然而,鱼类如何以及为什么保持稳定的集群形成仍然不清楚。本文采用多智能体深度强化学习和浸入边界格子Boltzmann方法的混合方法,对两个自由泳者的密集集群进行了数值研究。主动控制策略是通过同步训练领导者以给定的速度和方向游泳和跟随者保持靠近领导者来开发的。经过训练后,游泳运动员只有依靠尾拍的拍打才能抵抗强大的水动力,保持稳定的队形,同时沿预定的路线游泳。在稳定队形中,运动员的尾部运动是不规则和不对称的,表明运动员正在仔细地调整他们的身体运动学以平衡水动力。此外,随动者的平均振幅和运输成本显著降低,表明这些游泳运动员可以用更少的努力保持游泳速度。结果还表明,并排形成的流体动力学更稳定,但能量效率低于其他配置,而整个身体交错形成的整体能量效率更高。
Fish are observed to school in different configurations. However, how and why fish maintain a stable schooling formation still remains unclear. This work presents a numerical study of the dense schooling of two free swimmers by a hybrid method of the multi-agent deep reinforcement learning and the immersed boundary-lattice Boltzmann method. Active control policies are developed by synchronously training the leader to swim at a given speed and orientation and the follower to hold close proximity to the leader. After training, the swimmers could resist the strong hydrodynamic force to remain in stable formations and meantime swim in desired path, only by their tail-beat flapping. The tail movement of the swimmers in the stable formations are irregular and asymmetrical, indicating the swimmers are carefully adjusting their body-kinematics to balance the hydrodynamic force. In addition, a significant decrease in the mean amplitude and the cost of transport is found for the followers, indicating these swimmers could maintain the swimming speed with less efforts. The results also show that the side-by-side formation is hydrodynamically more stable but energetically less efficient than other configurations, while the full-body staggered formation is energetically more efficient as a whole.