Cooperative Data Collection With Multiple UAVs for Information Freshness in the Internet of Things

Cooperative Data Collection With Multiple UAVs for Information Freshness in the Internet of Things
复制标题

DOI:
10.1109/tcomm.2023.3255240
复制
发表时间:
2023-03
影响因子:
8.3
通讯作者:
Xijun Wang;Mengjie Yi;Juan Liu;Yan Zhang;Meng Wang;Bo Bai
Xijun Wang;Mengjie Yi;Juan Liu;Yan Zhang;Meng Wang;Bo Bai
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xijun Wang;Mengjie Yi;Juan Liu;Yan Zhang;Meng Wang;Bo Bai

文献摘要

相似文献

在物联网(IoT)中保持信息的新鲜度是一个关键但具有挑战性的问题。在本文中,我们研究了合作数据收集使用多个无人机(UAV)的目标是最小化的总平均信息年龄(AoI)。我们考虑各种约束的无人机,包括运动学,能量,轨迹,和碰撞避免,以优化数据收集过程。具体地说,每个无人机,这具有有限的机载能源,起飞从其初始位置和飞越传感器节点收集更新数据包与其他无人机合作。无人机必须在指定的时间段后以非负剩余能量降落在最终目的地,以确保它们有足够的能量完成任务。为了提高信息的新鲜度,设计无人机的飞行轨迹和传感器节点的传输调度至关重要。我们将多无人机数据收集问题建模为分散式部分可观察马尔可夫决策过程(Dec-POMDP),因为每个无人机都不知道环境的动态,只能观察一部分传感器。为了解决这个问题的挑战,我们提出了一种基于多智能体深度强化学习(DRL)的算法,具有集中式学习和分散式执行。除了奖励整形之外,我们还使用动作掩码来过滤掉无效的动作,并确保满足约束。仿真结果表明,与基线算法相比,所提算法能显著降低总平均AoI,且动作掩模方法的使用能提高算法的收敛速度。
Maintaining the freshness of information in the Internet of Things (IoT) is a critical yet challenging problem. In this paper, we study cooperative data collection using multiple Unmanned Aerial Vehicles (UAVs) with the objective of minimizing the total average Age of Information (AoI). We consider various constraints of the UAVs, including kinematic, energy, trajectory, and collision avoidance, in order to optimize the data collection process. Specifically, each UAV, which has limited on-board energy, takes off from its initial location and flies over sensor nodes to collect update packets in cooperation with the other UAVs. The UAVs must land at their final destinations with non-negative residual energy after the specified time duration to ensure they have enough energy to complete their missions. It is crucial to design the trajectories of the UAVs and the transmission scheduling of the sensor nodes to enhance information freshness. We model the multi-UAV data collection problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), as each UAV is unaware of the dynamics of the environment and can only observe a part of the sensors. To address the challenges of this problem, we propose a multi-agent Deep Reinforcement Learning (DRL)-based algorithm with centralized learning and decentralized execution. In addition to the reward shaping, we use action masks to filter out invalid actions and ensure that the constraints are met. Simulation results demonstrate that the proposed algorithms can significantly reduce the total average AoI compared to the baseline algorithms, and the use of the action mask method can improve the convergence speed of the proposed algorithm.