Deep Reinforcement Learning Driven UAV-Assisted Edge Computing

Deep Reinforcement Learning Driven UAV-Assisted Edge Computing
复制标题

深度强化学习驱动的无人机辅助边缘计算

DOI:
10.1109/jiot.2022.3196842
复制
发表时间:
2022-12
影响因子:
10.6
通讯作者:
Liang Zhang;B. Jabbari;Nirwan Ansari
Liang Zhang;B. Jabbari;Nirwan Ansari
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liang Zhang;B. Jabbari;Nirwan Ansari

文献摘要

相似文献

无人机(UAV)在满足物联网设备(IoT)的即时连接和计算需求方面发挥着关键作用,特别是在危机和灾难管理方面。在这项工作中,我们专注于优化无人机的轨迹沿着,物联网设备在多个时隙中提供通信和计算资源。IoTD的体验质量(QoE)取决于其延迟性能;因此,我们的目标是最大限度地提高所有IoTD在整个时隙中的平均聚合QoE。然而,这是一个非凸、非线性、混合离散的优化问题,很难求解并得到最优解。因此,我们提出了两种深度强化学习算法来解决这个问题,考虑无人机路径规划,用户分配,带宽和计算资源分配。我们比较了我们提出的算法的性能,通过模拟与三个基线情况:1)与固定的无人机位置; 2)无无人机;和3)固定的无人机轨迹。我们证明了深度强化学习算法的性能优于所有基线情况。
Unmanned aerial vehicles (UAVs) are playing a critical role in provisioning instant connectivity and computational needs of Internet of Things Devices (IoTDs), especially in crisis and disaster management. In this work, we focus on optimizing trajectories of UAVs along which IoTDs are served with communication and computing resources in multiple time slots. The Quality of Experience (QoE) of an IoTD depends on its latency performance; we thus aim to maximize the average aggregate QoE of all IoTDs overall time slots. However, this is a nonconvex, nonlinear, and mixed discrete optimization problem, which is difficult to solve and obtain the optimal solution. We thus propose two deep reinforcement learning algorithms to solve this problem by considering UAV path planning, user assignment, bandwidth, and computing resource assignment. We compare the performance of our proposed algorithms through simulations with three baseline cases: 1) with fixed UAV locations; 2) without UAVs; and 3) the fixed UAV trajectories. We demonstrate that the deep reinforcement learning algorithms perform better than all baseline cases.