Fast Reinforcement Learning for Energy-Efficient Wireless Communication

Fast Reinforcement Learning for Energy-Efficient Wireless Communication
复制标题

DOI:
10.1109/tsp.2011.2165211
复制
发表时间:
2010-09
影响因子:
5.4
通讯作者:
Nicholas Mastronarde;M. Schaar
Nicholas Mastronarde;M. Schaar
中科院分区:
工程技术1区
文献类型:
--
作者:
Nicholas Mastronarde;M. Schaar

文献摘要

被引文献

相似文献

我们考虑了延迟敏感数据(如多媒体数据)在衰落信道上的节能点对点传输问题。我们提出了一个严格和统一的框架,同时利用物理层和系统级技术,在随机和未知的交通和信道条件下,在延迟约束下最大限度地减少能量消耗。我们将问题表述为一个马尔可夫决策过程,并使用强化学习在线解决它。所提出的在线方法的优点是:i)它不需要先验的流量到达和信道统计信息来确定最优的物理层和系统级电源管理策略;Ii)它利用了系统的部分信息,因此需要学习的信息比使用传统强化学习算法时要少;iii)避免了对动作探索的需要,这严重限制了传统强化学习算法的自适应速度和运行时性能。
We consider the problem of energy-efficient point-to-point transmission of delay-sensitive data (e.g., multimedia data) over a fading channel. We propose a rigorous and unified framework for simultaneously utilizing both physical-layer and system-level techniques to minimize energy consumption, under delay constraints, in the presence of stochastic and unknown traffic and channel conditions. We formulate the problem as a Markov decision process and solve it online using reinforcement learning. The advantages of the proposed online method are that i) it does not require a priori knowledge of the traffic arrival and channel statistics to determine the jointly optimal physical-layer and system-level power management strategies; ii) it exploits partial information about the system so that less information needs to be learned than when using conventional reinforcement learning algorithms; and iii) it obviates the need for action exploration, which severely limits the adaptation speed and run-time performance of conventional reinforcement learning algorithms.