TACT: A Transfer Actor-Critic Learning Framework for Energy Saving in Cellular Radio Access Networks

TACT: A Transfer Actor-Critic Learning Framework for Energy Saving in Cellular Radio Access Networks
复制标题

DOI:
10.1109/twc.2014.022014.130840
复制
发表时间:
2014-04-01
影响因子:
10.4
通讯作者:
Zhang, Honggang
Zhang, Honggang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Rongpeng;Zhao, Zhifeng;Zhang, Honggang

文献摘要

被引文献

相似文献

最近的工作已经验证了通过动态地打开/关闭一些基站(BS)来实现的提高无线电接入网络(RAN)中的能量效率的可能性。在本文中,我们扩展了BS交换操作,这应该与流量负载的变化相匹配的研究。而不是依赖于动态的流量负载,这仍然是相当具有挑战性的精确预测,我们首先制定的流量变化作为一个马尔可夫决策过程。之后,为了前瞻性地最小化RAN的能量消耗,我们设计了一个基于强化学习框架的BS切换操作方案。此外,为了加快正在进行的学习过程中,转移演员-评论家算法(TACT),它利用历史时期或相邻区域的转移学习经验,提出并证明收敛。最后,我们评估我们提出的方案,在各种实际配置下进行了广泛的模拟,并表明,建议的TACT算法有助于性能的跳跃启动,并证明了在可容忍的延迟性能为代价的显着的能源效率提高的可行性。
Recent works have validated the possibility of improving energy efficiency in radio access networks (RANs), achieved by dynamically turning on/off some base stations (BSs). In this paper, we extend the research over BS switching operations, which should match up with traffic load variations. Instead of depending on the dynamic traffic loads which are still quite challenging to precisely forecast, we firstly formulate the traffic variations as a Markov decision process. Afterwards, in order to foresightedly minimize the energy consumption of RANs, we design a reinforcement learning framework based BS switching operation scheme. Furthermore, to speed up the ongoing learning process, a transfer actor-critic algorithm (TACT), which utilizes the transferred learning expertise in historical periods or neighboring regions, is proposed and provably converges. In the end, we evaluate our proposed scheme by extensive simulations under various practical configurations and show that the proposed TACT algorithm contributes to a performance jumpstart and demonstrates the feasibility of significant energy efficiency improvement at the expense of tolerable delay performance.