Online Learning for Position-Aided Millimeter Wave Beam Training

Online Learning for Position-Aided Millimeter Wave Beam Training
复制标题

DOI:
10.1109/access.2019.2902372
复制
发表时间:
2018-09
期刊:
影响因子:
3.9
通讯作者:
Vutha Va;Takayuki Shimizu;G. Bansal;R. Heath
Vutha Va;Takayuki Shimizu;G. Bansal;R. Heath
中科院分区:
计算机科学3区
文献类型:
--
作者:
Vutha Va;Takayuki Shimizu;G. Bansal;R. Heath

文献摘要

被引文献

相似文献

精确的波束对准对于基于波束的毫米波通信至关重要。传统的波束扫描解决方案通常开销较大,这对于诸如车与万物通信之类的移动应用来说是不可接受的。基于学习的解决方案利用传感器数据(例如位置信息)来确定良好的波束方向,这是一种降低开销的方法。然而,现有的大多数解决方案都是监督学习,其训练数据是预先收集的。在本文中,我们采用多臂老虎机框架来开发用于波束对选择和优化的在线学习算法。波束对选择算法在某些预定义的波束码本中学习粗略的波束方向,例如以离散角度,间隔为3分贝波束宽度。波束优化则对已确定的方向进行微调,以匹配该位置处功率角谱的峰值。波束对选择采用带有新提出的风险感知特征的上置信界,而波束优化采用改进的乐观优化算法。所提出的算法能够快速学习推荐良好的波束对。当发射端和接收端均使用16×16阵列时,在未优化的码本上,仅用30个波束对的训练预算,在100个时间步内,与穷举搜索(遍历271×271个波束对)相比,平均可实现1分贝的增益。
Accurate beam alignment is essential for the beam-based millimeter wave communications. The conventional beam sweeping solutions often have large overhead, which is unacceptable for mobile applications, such as a vehicle to everything. The learning-based solutions that leverage the sensor data (e.g., position) to identify the good beam directions are one approach to reduce the overhead. Most existing solutions, though, are supervised learning, where the training data are collected beforehand. In this paper, we use a multi-armed bandit framework to develop the online learning algorithms for beam pair selection and refinement. The beam pair selection algorithm learns coarse beam directions in some predefined beam codebook, e.g., in discrete angles, separated by the 3 dB beamwidths. The beam refinement fine-tunes the identified directions to match the peak of the power angular spectrum at that position. The beam pair selection uses the upper confidence bound with a newly proposed risk-aware feature, while the beam refinement uses a modified optimistic optimization algorithm. The proposed algorithms learn to recommend the good beam pairs quickly. When using $16\times 16$ arrays at both transmitter and receiver, it can achieve, on average, 1-dB gain over the exhaustive search (over $271\times 271$ beam pairs) on the unrefined codebook within 100 time steps with a training budget of only 30 beam pairs.