Reinforcement Learning-Based Adaptive Transmission in Time-Varying Underwater Acoustic Channels

Reinforcement Learning-Based Adaptive Transmission in Time-Varying Underwater Acoustic Channels
复制标题

DOI:
10.1109/access.2017.2784239
复制
发表时间:
2018
期刊:
影响因子:
3.9
通讯作者:
Chaofeng Wang;Zhaohui Wang;Wensheng Sun;D. Fuhrmann
Chaofeng Wang;Zhaohui Wang;Wensheng Sun;D. Fuhrmann
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chaofeng Wang;Zhaohui Wang;Wensheng Sun;D. Fuhrmann

文献摘要

被引文献

相似文献

本文研究了在一个长时间逐历元工作的水声点对点通信系统中的自适应传输问题。固定量的信息比特周期性地到达发射机数据队列,并且等待经由每个时期内的多个分组的传输。为了权衡能量消耗与传输延迟,发射机基于数据队列状态和当前和未来时期的预测信道条件来决定每个时期开始时的传输动作,包括发送或不发送,以及发送功率和调制编码参数。为了描述水声信道的快速衰落和大规模阴影,每个历元内的信道的特征在于由复合Nakagami-lognormal分布,和分布参数的演变建模为一个未知的马尔可夫过程。鉴于信道只能在主动传输过程中被观察到,我们将自适应传输问题表示为部分可观察的马尔可夫决策过程,并在基于模型的强化学习框架中开发了一种在线算法。该算法递归地估计信道模型参数,跟踪信道动态,并计算使长期系统成本最小化的最佳传输动作。仿真结果的基础上,从两个领域的实验信道测量表明,该算法实现体面的性能相对于基准方法,假设完美的和非因果的信道知识。
This paper studies adaptive transmission in an underwater acoustic (UWA) point-to-point communication system that operates on an epoch-by-epoch basis for a long term. A fixed amount of information bits periodically arrive at the transmitter data queue, and wait for transmission via a number of packets within each epoch. To trade off energy consumption with transmission latency, the transmitter decides the transmission action at the beginning of each epoch, including to transmit or not, and the transmission power and the modulation-and-coding parameters, based on the data queue status and the predicted channel conditions in the current and future epochs. To describe both the fast fading and the large-scale shadowing of UWA channels, the channel within each epoch is characterized by a compound Nakagami-lognormal distribution, and the evolution of the distribution parameters is modeled as an unknown Markov process. Given that the channel can only be observed during active transmissions, we formulate the adaptive transmission problem as a partially observable Markov decision process, and develop an online algorithm in a model-based reinforcement learning framework. The algorithm recursively estimates the channel model parameters, tracks the channel dynamics, and computes the optimal transmission action that minimizes a long-term system cost. Emulated results based on channel measurements from two-field experiments demonstrate that the proposed algorithm achieves decent performance relative to a benchmark method that assumes perfect and non-causal channel knowledge.