Multi-Agent Deep Reinforcement Learning for Spectral Efficiency Optimization in Vehicular Optical Camera Communications

Multi-Agent Deep Reinforcement Learning for Spectral Efficiency Optimization in Vehicular Optical Camera Communications
复制标题

DOI:
10.1109/tmc.2023.3278277
复制
发表时间:
2024-05
影响因子:
7.9
通讯作者:
Amirul Islam;N. Thomos;Leila Musavian
Amirul Islam;N. Thomos;Leila Musavian
中科院分区:
计算机科学2区
文献类型:
--
作者:
Amirul Islam;N. Thomos;Leila Musavian

文献摘要

相似文献

在这篇文章中,我们提出了一种车载光学摄像机通信系统,可以满足低误码率(BER)和超低延迟的限制。首先,我们制定了一个总和频谱效率优化问题,旨在找到车辆的速度和调制阶数,最大限度地提高总和频谱效率的可靠性和延迟约束。该问题是一个具有非线性约束的混合整数规划问题,即使对于一个小的调制阶数集,也是NP-难的。为了克服所带来的高计算和时间复杂性,防止其与传统方法的解决方案,我们首先建模的优化问题作为一个部分可观察的马尔可夫决策过程。然后,我们使用独立的Q学习框架来解决它,其中每个车辆都充当独立的代理。由于状态-动作空间很大,因此我们采用深度强化学习(DRL)来有效地解决它。由于问题的约束,我们采用拉格朗日松弛的方法之前,解决它使用DRL框架。仿真结果表明,所提出的基于DRL的优化方案可以有效地学习如何在满足BER和超低延迟约束的情况下最大化总频谱效率。评估进一步表明,我们的计划可以实现上级性能相比,基于无线电频率的车载通信系统和其他车辆OCC的变种,我们的计划。
In this article, we propose a vehicular optical camera communication system that can meet low bit error rate (BER) and ultra-low latency constraints. First, we formulate a sum spectral efficiency optimization problem that aims at finding the speed of vehicles and the modulation order that maximizes the sum spectral efficiency subject to reliability and latency constraints. This problem is mixed-integer programming with nonlinear constraints, and even for a small set of modulation orders, is NP-hard. To overcome the entailed high computational and time complexity which prevents its solution with traditional methods, we first model the optimization problem as a partially observable Markov decision process. We then solve it using an independent Q-learning framework, where each vehicle acts as an independent agent. Since the state-action space is large we then adopt deep reinforcement learning (DRL) to solve it efficiently. As the problem is constrained, we employ the Lagrange relaxation approach prior to solving it using the DRL framework. Simulation results demonstrate that the proposed DRL-based optimization scheme can effectively learn how to maximize the sum spectral efficiency while satisfying the BER and ultra-low latency constraints. The evaluation further shows that our scheme can achieve superior performance compared to radio frequency-based vehicular communication systems and other vehicular OCC variants of our scheme.