Wireless Resource Scheduling in Virtualized Radio Access Networks Using Stochastic Learning

Wireless Resource Scheduling in Virtualized Radio Access Networks Using Stochastic Learning
复制标题

DOI:
10.1109/tmc.2017.2742949
复制
发表时间:
2018-04
影响因子:
7.9
通讯作者:
Xianfu Chen;Zhu Han;Honggang Zhang;G. Xue;Yong Xiao;M. Bennis
Xianfu Chen;Zhu Han;Honggang Zhang;G. Xue;Yong Xiao;M. Bennis
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xianfu Chen;Zhu Han;Honggang Zhang;G. Xue;Yong Xiao;M. Bennis

文献摘要

被引文献

相似文献

在密集无线接入网络(RAN)中如何分配有限的无线资源仍然具有挑战性。通过利用软件定义的控制平面,独立的基站(BS)被虚拟化为一个集中式网络控制器(CNC)。这种虚拟化将CNC与无线服务提供商(WSP)解耦。我们研究一种虚拟RAN,其中CNC在调度时隙开始时根据其订阅的WSP的出价向移动终端(MT)拍卖信道。每个WSP旨在最大化从出价信道中获得的预期长期收益,以满足MT传输数据包的需求。我们将该问题表述为一个随机博弈,其中一个WSP的信道拍卖和分组调度决策取决于网络状态及其竞争对手的控制策略。为了接近均衡解,提出了一种具有有界遗憾的抽象随机博弈。每个WSP的决策过程被建模为一个马尔可夫决策过程(MDP)。为了解决信令开销和计算复杂性问题,我们将MDP分解为一系列具有缩减状态空间的单智能体MDP,并推导出一种在线局部化算法来学习状态值函数。我们的结果表明,在每个MT的平均效用方面有显著的性能提升。
How to allocate the limited wireless resource in dense radio access networks (RANs) remains challenging. By leveraging a software-defined control plane, the independent base stations (BSs) are virtualized as a centralized network controller (CNC). Such virtualization decouples the CNC from the wireless service providers (WSPs). We investigate a virtualized RAN, where the CNC auctions channels at the beginning of scheduling slots to the mobile terminals (MTs) based on bids from their subscribing WSPs. Each WSP aims at maximizing the expected long-term payoff from bidding channels to satisfy the MTs for transmitting packets. We formulate the problem as a stochastic game, where the channel auction and packet scheduling decisions of a WSP depend on the state of network and the control policies of its competitors. To approach the equilibrium solution, an abstract stochastic game is proposed with bounded regret. The decision making process of each WSP is modeled as a Markov decision process (MDP). To address the signalling overhead and computational complexity issues, we decompose the MDP into a series of single-agent MDPs with reduced state spaces, and derive an online localized algorithm to learn the state value functions. Our results show significant performance improvements in terms of per-MT average utility.