Spectrum Sharing in Vehicular Networks Based on Multi-Agent Reinforcement Learning

Spectrum Sharing in Vehicular Networks Based on Multi-Agent Reinforcement Learning
复制标题

DOI:
10.1109/jsac.2019.2933962
复制
发表时间:
2019-10-01
影响因子:
16.4
通讯作者:
Li, Geoffrey Ye
Li, Geoffrey Ye
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liang, Le;Ye, Hao;Li, Geoffrey Ye

文献摘要

被引文献

相似文献

研究了基于多智能体强化学习的车载网络频谱共享问题,其中多个车到车(V2V)链路重复使用被车到基础设施(V2I)链路占用的频谱。高机动性车辆环境中的快速信道变化排除了在基站收集准确的瞬时信道状态信息用于集中资源管理的可能性。作为响应,我们将资源共享建模为多智能体强化学习问题,然后使用适合于分布式实现的基于指纹的深度Q网络方法来解决该问题。V2V链路,每个都充当代理,共同与通信环境交互,接收不同的观测结果,但获得共同的奖励,并学习通过使用获得的经验更新Q网络来改进频谱和功率分配。我们证明,在适当的奖励设计和培训机制下,多个V2V代理成功地学习以分布式方式合作,同时提高V2I链路的总和容量和V2V链路的有效载荷投递率。
This paper investigates the spectrum sharing problem in vehicular networks based on multi-agent reinforcement learning, where multiple vehicle-to-vehicle (V2V) links reuse the frequency spectrum preoccupied by vehicleto-infrastructure (V2I) links. Fast channel variations in high mobility vehicular environments preclude the possibility of collecting accurate instantaneous channel state information at the base station for centralized resource management. In response, we model the resource sharing as a multi-agent reinforcement learning problem, which is then solved using a fingerprint-based deep Q-network method that is amenable to a distributed implementation. The V2V links, each acting as an agent, collectively interact with the communication environment, receive distinctive observations yet a common reward, and learn to improve spectrum and power allocation through updating Q-networks using the gained experiences. We demonstrate that with a proper reward design and training mechanism, the multiple V2V agents successfully learn to cooperate in a distributed way to simultaneously improve the sum capacity of V2I links and payload delivery rate of V2V links.