Contextual Bandit Learning for Machine Type Communications in the Null Space of Multi-Antenna Systems

Contextual Bandit Learning for Machine Type Communications in the Null Space of Multi-Antenna Systems
复制标题

DOI:
10.1109/tcomm.2019.2955454
复制
发表时间:
2020-02-01
影响因子:
8.3
通讯作者:
Haapola, Jussi
Haapola, Jussi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ali, Samad;Asgharimoghaddam, Hossein;Haapola, Jussi

文献摘要

被引文献

相似文献

由于机器类型通信(MTC)对蜂窝用户的干扰,确保传统宽带蜂窝用户与机器类型通信(MTCS)的有效共存是具有挑战性的。该干扰挑战源于这样一个事实,即由于MTC业务的小分组性质,从机器类型设备(MTD)到蜂窝基站(BS)的信道状态信息(CSI)的获取是不可行的。本文提出了一种基于机会空间正交化(OSO)概念的MTC与传统蜂窝通信干扰管理的新方法。具体地,考虑具有多天线BS和机器类型聚合器(MTA)的蜂窝系统,其中,接收波束形成器被设计为使蜂窝用户的速率最大化,机器类型聚合器(MTA)从大量MTD接收数据。BS和MTA共享相同的上行链路资源,因此,MTD传输在BS上产生干扰。然而,如果在每个波束形成器的每个给定时间有大量MTD可供选择用于传输,则可以选择一个MTD,使得它几乎不会对BS造成干扰。对多个MTDS对同一波束形成器的干扰特性进行了全面的分析研究。证明了对于每个波束形成器,MTD的存在使得对BS的干扰可以忽略不计。为了进一步研究这种干扰,推导了蜂窝用户的信干噪比的分布,并在此基础上给出了中断概率的分布。然而,OSO的优化实现需要BS中所有链路的CSI,这对于MTC是不现实的。针对这一问题,提出了一种基于上下文多武装匪徒(MAB)学习概念的在线学习方法。接收波束形成器用作上下文MAB设置和汤普森采样的上下文:提出了一种解决上下文MAB问题的众所周知的方法。由于该设置中的上下文数目可以是无限的,因此需要近似汤普森抽样的后验分布。对于给定的波束形成器,提出了两种函数逼近方法,a)线性全后验抽样,b)神经网络,用于最佳选择传输的MTD。仿真结果表明,在MTDS到BS之间不存在CSI的情况下实现OSO是可能的。当所有MTDS到BS的CSI已知时,线性全后验抽样可获得近90%的最优分配。
Ensuring an effective coexistence of conventional broadband cellular users with machine type communications (MTCs) is challenging due to the interference from MTCs to cellular users. This interference challenge stems from the fact that the acquisition of channel state information (CSI) from machine type devices (MTD) to cellular base stations (BS) is infeasible due to the small packet nature of MTC traffic. In this paper, a novel approach based on the concept of opportunistic spatial orthogonalization (OSO) is proposed for interference management between MTC and conventional cellular communications. In particular, a cellular system is considered with a multi-antenna BS in which a receive beamformer is designed to maximize the rate of a cellular user, and, a machine type aggregator (MTA) that receives data from a large set of MTDs. The BS and MTA share the same uplink resources, and, therefore, MTD transmissions create interference on the BS. However, if there is a large number of MTDs to chose from for transmission at each given time for each beamformer, one MTD can be selected such that it causes almost no interference on the BS. A comprehensive analytical study of the characteristics of such an interference from several MTDs on the same beamformer is carried out. It is proven that, for each beamformer, an MTD exists such that the interference on the BS is negligible. To further investigate such interference, the distribution of the signal-to-interference-plus-noise ratio (SINR) of the cellular user is derived, and, subsequently, the distribution of the outage probability is presented. However, the optimal implementation of OSO requires the CSI of all the links in the BS, which is not practical for MTC. To solve this problem, an online learning method based on the concept of contextual multi-armed bandits (MAB) learning is proposed. The receive beamformer is used as the context of the contextual MAB setting and Thompson sampling: a well-known method of solving contextual MAB problems is proposed. Since the number of contexts in this setting can be unlimited, approximating the posterior distributions of Thompson sampling is required. Two function approximation methods, a) linear full posterior sampling, and, b) neural networks are proposed for optimal selection of MTD for transmission for the given beamformer. Simulation results show that is possible to implement OSO with no CSI from MTDs to the BS. Linear full posterior sampling achieves almost 90% of the optimal allocation when the CSI from all the MTDs to the BS is known.