Reinforcement Learning-based Trajectory Optimization for Data Muling with Underwater Mobile Nodes

Reinforcement Learning-based Trajectory Optimization for Data Muling with Underwater Mobile Nodes
复制标题

DOI:
10.1109/access.2022.3165046
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Qiang Fu;A. Song;Fumin Zhang;Miao Pan
Qiang Fu;A. Song;Fumin Zhang;Miao Pan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Qiang Fu;A. Song;Fumin Zhang;Miao Pan

文献摘要

相似文献

本文研究了具有移动的节点的水下数据传输轨迹优化问题。在水下数据采集场景中,多个自主水下航行器(AUV)探索或采样使命区域,自主水面航行器(ASV)访问正在进行的AUV以检索收集的数据。优化目标是同时最大化数据传输的公平性和最小化表面节点的旅行距离。我们提出了一个最近K强化学习算法。在该算法中,我们只选择最近的K个AUV作为下一个节点的数据传输的候选人。我们选择AUV和ASV之间的距离作为状态,选择AUV作为动作。奖励被设计为传输的数据量和ASV行驶距离的函数。在多个ASV的场景中,AUV的关联策略,提出了支持使用多个表面节点。我们进行计算机模拟性能评估。研究了AUV数量、使命区域大小和状态选择对系统性能的影响。仿真结果表明,该算法在公平性和ASV行驶距离方面优于传统方法。
This manuscript addresses the trajectory optimization for underwater data muling with mobile nodes. In the underwater data muling scenario, multiple autonomous underwater vehicles (AUVs) explore or sample a mission area and autonomous surface vehicles (ASVs) visit underway AUVs to retrieve collected data. The optimization objectives are to simultaneously maximize fairness in data transmissions and minimize the travel distance of the surface nodes. We propose a nearest-K reinforcement learning algorithm. In the algorithm, we choose only from the nearest-K AUVs as candidates for the next node for data transmissions. We choose the distance between AUVs and the ASV as the state, selected AUVs as the action. A reward is designed as the function of both data volume transmitted and the ASV travel distance. In the scenario with multiple ASVs, an AUV association strategy is proposed to support the use of multiple surface nodes. We conduct computer simulations for performance evaluation. The effects from the number of AUVs, the size of the mission area, and state selection are investigated. Simulation results show that the proposed algorithm outperforms traditional methods in terms of fairness and the ASV travel distance.