Wireless Access Control in Edge-Aided Disaster Response: A Deep Reinforcement Learning-Based Approach

Wireless Access Control in Edge-Aided Disaster Response: A Deep Reinforcement Learning-Based Approach
复制标题

DOI:
10.1109/access.2021.3067662
复制
发表时间:
2021
期刊:
影响因子:
3.9
通讯作者:
Hang Zhou;Xiaoyan Wang;M. Umehira;Xianfu Chen;Celimuge Wu;Yusheng Ji
Hang Zhou;Xiaoyan Wang;M. Umehira;Xianfu Chen;Celimuge Wu;Yusheng Ji
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hang Zhou;Xiaoyan Wang;M. Umehira;Xianfu Chen;Celimuge Wu;Yusheng Ji

文献摘要

被引文献

相似文献

在重大灾害发生后,通信基础设施最有可能遭到破坏,这将导致灾区进一步混乱。现代救援活动在很大程度上依赖于无线通信,例如安全状态报告、中断区域监测、疏散指示、救援协调等。即使在正常通信基础设施退化或被破坏的情况下,也必须以快速可靠的方式传输和处理从受害者、传感器和响应者产生的大量数据。为此,通过在边缘部署MDRU(可移动和可部署资源单元)和中继单元来重建灾后网络是一个非常有前途的解决方案。然而,由于灾后环境的频繁变化和缺乏事先统计信息,在这种匆忙形成的异构网络中进行最优无线接入控制极具挑战性。在本文中,我们提出了一种基于学习的边缘辅助灾难响应网络的无线访问控制方法。更具体地说,我们将无线接入控制过程建模为离散时间单代理马尔可夫决策过程,并利用深度强化学习技术解决问题。通过大量的仿真结果,我们表明,所提出的机制显着优于基线方案的延迟和丢包率。
The communication infrastructure is most likely to be damaged after a major disaster occurred, which would lead to further chaos in the disaster stricken area. Modern rescue activities heavily rely on the wireless communications, such as safety status report, disrupted area monitoring, evacuation instruction, rescue coordination, etc. Large amount of data generated from victims, sensors and responders must be delivered and processed in a fast and reliable way, even when the normal communication infrastructure is degraded or destroyed. To this end, reconstructing the post-disaster network by deploying MDRU (Movable and Deployable Resource Unit) and relay unit at edge is a very promising solution. However, the optimal wireless access control in this heterogeneous hastily formed network is extremely challenging, due to the frequent varying environment and the lack of statistics information in advance in post-disaster scenarios. In this paper, we propose a learning based wireless access control approach for edge-aided disaster response network. More specifically, we model the wireless access control procedure as a discrete-time single agent Markov decision process, and solve the problem by exploiting deep reinforcement learning technique. By extensive simulation results, we show that the proposed mechanism significantly outperforms the baseline schemes in terms of delay and packet drop rate.