Deep Multi-User Reinforcement Learning for Distributed Dynamic Spectrum Access

Deep Multi-User Reinforcement Learning for Distributed Dynamic Spectrum Access
复制标题

DOI:
10.1109/twc.2018.2879433
复制
发表时间:
2019-01-01
影响因子:
10.4
通讯作者:
Cohen, Kobi
Cohen, Kobi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Naparstek, Oshri;Cohen, Kobi

文献摘要

被引文献

相似文献

研究了多信道无线网络中网络效用最大化的动态频谱接入问题。共享带宽被划分为K个正交信道。在每个时隙的开始,每个用户选择一个信道并以一定的传输概率发送一个分组。在每个时隙之后,已经发送分组的每个用户接收指示其分组是否被成功递送的本地观测(即,ACK信号)。我们的目标是一个多用户的频谱访问策略,最大限度地提高了一定的网络效用,在分布式的方式,没有在线协调或用户之间的消息交换。一般来说,由于状态空间大且状态的部分可观测性,获得频谱接入问题的最优解在计算上是昂贵的。针对这一问题,提出了一种基于深度多用户强化学习的分布式动态频谱接入算法。具体地,在每个时隙,每个用户基于用于最大化目标函数的训练的深度Q网络将其当前状态映射到频谱接入动作。博弈论分析的系统动力学建立设计原则的算法的实施。实验结果表明,该算法具有较好的性能.
We consider the problem of dynamic spectrum access for network utility maximization in multichannel wireless networks. The shared bandwidth is divided into K orthogonal channels. In the beginning of each time slot, each user selects a channel and transmits a packet with a certain transmission probability. After each time slot, each user that has transmitted a packet receives a local observation indicating whether its packet was successfully delivered or not (i.e., ACK signal). The objective is a multi-user strategy for accessing the spectrum that maximizes a certain network utility in a distributed manner without online coordination or message exchanges between users. Obtaining an optimal solution for the spectrum access problem is computationally expensive, in general, due to the large-state space and partial observability of the states. To tackle this problem, we develop a novel distributed dynamic spectrum access algorithm based on deep multi-user reinforcement leaning. Specifically, at each time slot, each user maps its current state to the spectrum access actions based on a trained deep-Q network used to maximize the objective function. Game theoretic analysis of the system dynamics is developed for establishing design principles for the implementation of the algorithm. The experimental results demonstrate the strong performance of the algorithm.