Contextual-Bandit based MIMO Relay Selection Policy with Channel Uncertainty

Contextual-Bandit based MIMO Relay Selection Policy with Channel Uncertainty
复制标题

DOI:
10.1109/icc40277.2020.9148879
复制
发表时间:
2020-06
期刊:
ICC 2020 - 2020 IEEE International Conference on Communications (ICC)
影响因子:
--
通讯作者:
Ankit Gupta;Naveen Mysore Balasubramanya;M. Sellathurai
Ankit Gupta;Naveen Mysore Balasubramanya;M. Sellathurai
中科院分区:
其他
文献类型:
--
作者:
Ankit Gupta;Naveen Mysore Balasubramanya;M. Sellathurai

文献摘要

被引文献

相似文献

在这项工作中,我们挖掘了多臂盗贼方案在协作多输入多输出(MIMO)无线网络中的潜在好处。特别地,我们考虑了一种在线放大转发MIMO中继选择(RS)策略,其中为中继提供了不确定的信道状态信息(CSI)。本文将RS策略设计为一种序贯经验驱动的学习算法,该算法使用上下文向量提供的不完美CSI和当前策略获得的奖励的过去经验来学习选择最优的中继节点,目标是最大化累积平均奖励。此外,通过大量的仿真结果,我们证明了所提出的基于CB的RS策略比传统的Gram-Schmidt方法获得了更好的性能提升。
In this work, we exploit the potential benefits of multi-arm bandit scheme in cooperative multiple-input multiple output (MIMO) wireless networks. In particular, we consider an online-policy for amplify-and-forward MIMO relay selection (RS), where relays are provided with uncertain channel state information (CSI). We design the RS policy as a sequential experience-driven learning algorithm with a contextual bandit (CB) approach, where the algorithm learns to select an optimal relay node using the imperfect CSI provided as a context vector and the past experience of rewards procured with current policy, with the aim of maximizing the cumulative mean reward over time. Further, with extensive simulation result, we demonstrate that proposed CB based RS policy achieves superior performance gains compared to conventional Gram-Schmidt method.