On-Demand Channel Bonding in Heterogeneous WLANs: A Multi-Agent Deep Reinforcement Learning Approach

On-Demand Channel Bonding in Heterogeneous WLANs: A Multi-Agent Deep Reinforcement Learning Approach
复制标题

DOI:
10.3390/s20102789
复制
发表时间:
2020-05
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
H. Qi;Hao Huang;Zhiqun Hu;X. Wen;Zhaoming Lu
H. Qi;Hao Huang;Zhiqun Hu;X. Wen;Zhaoming Lu
中科院分区:
其他
文献类型:
--
作者:
H. Qi;Hao Huang;Zhiqun Hu;X. Wen;Zhaoming Lu

文献摘要

相似文献

为了满足无线局域网(WLAN)不断增长的流量需求,IEEE 802.11标准引入了信道绑定。虽然信道绑定有效地提高了传输速率,但较宽的信道减少了非重叠信道的数量,并且更容易受到干扰。同时,流量负载从一个接入点(AP)到另一个接入点是不同的,并且根据一天中的时间而显著变化。因此,应仔细选择主信道和信道绑定带宽,以满足业务需求并保证性能增益。在本文中,我们提出了一种基于深度强化学习(DRL)的按需信道绑定(O-DCB)算法,用于异构WLAN,以减少传输延迟,其中AP具有不同的信道绑定能力。在这个问题中,状态空间是连续的,动作空间是离散的。然而,行动空间的大小与AP的数量呈指数增长,使用单代理DRL,这严重影响了学习率。为了加速学习,多智能体深度确定性策略梯度(MADDPG)用于训练O-DCB。从校园无线局域网收集的真实的流量轨迹用于训练和测试O-DCB。仿真结果表明,该算法具有良好的收敛性和较低的延迟比其他算法。
In order to meet the ever-increasing traffic demand of Wireless Local Area Networks (WLANs), channel bonding is introduced in IEEE 802.11 standards. Although channel bonding effectively increases the transmission rate, the wider channel reduces the number of non-overlapping channels and is more susceptible to interference. Meanwhile, the traffic load differs from one access point (AP) to another and changes significantly depending on the time of day. Therefore, the primary channel and channel bonding bandwidth should be carefully selected to meet traffic demand and guarantee the performance gain. In this paper, we proposed an On-Demand Channel Bonding (O-DCB) algorithm based on Deep Reinforcement Learning (DRL) for heterogeneous WLANs to reduce transmission delay, where the APs have different channel bonding capabilities. In this problem, the state space is continuous and the action space is discrete. However, the size of action space increases exponentially with the number of APs by using single-agent DRL, which severely affects the learning rate. To accelerate learning, Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is used to train O-DCB. Real traffic traces collected from a campus WLAN are used to train and test O-DCB. Simulation results reveal that the proposed algorithm has good convergence and lower delay than other algorithms.