Look-Ahead and Learning Approaches for Energy Harvesting Communications Systems

Look-Ahead and Learning Approaches for Energy Harvesting Communications Systems
复制标题

DOI:
10.1109/tgcn.2019.2953644
复制
发表时间:
2020-03-01
影响因子:
4.8
通讯作者:
Kamal, Ahmed E.
Kamal, Ahmed E.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Masadeh, Ala'eddin;Wang, Zhengdao;Kamal, Ahmed E.

文献摘要

被引文献

相似文献

这项工作研究了能量收集通信系统的性能。该系统由一个发射机和一个接收机组成。发射机配备了一个无限的缓冲区来存储数据,以及能量收集能力,以收集可再生能源并将其存储在有限的电池中。目标是使这种系统的预期累积吞吐量最大化。将寻找最优功率分配策略的问题表示为马尔可夫决策过程。两种情况下被认为是基于信道增益和能量收集过程的统计知识的可用性。当该知识可用时,算法被设计为最大化预期吞吐量,同时降低传统方法的复杂性(例如,值迭代)。该算法利用关于信道、收获的能量和当前电池水平的即时知识来找到接近最优的策略。对于第二种情况,当统计知识不可用时,使用强化学习。使用两种不同的探索算法,基于收敛和ε贪婪算法。仿真和与传统算法的比较表明,当统计知识可用时,前瞻算法的有效性,以及当这些知识不可用时,强化学习在优化系统性能方面的有效性。
This work investigates the performance of an energy harvesting communications system. This system consists of a transmitter and a receiver. The transmitter is equipped with an infinite buffer to store data, and energy harvesting capability to harvest renewable energy and store it in a finite battery. The goal is to maximize the expected cumulative throughput of such systems. The problem of finding an optimal power allocation policy is formulated as a Markov decision process. Two cases are considered based on the availability of statistical knowledge about the channel gain and energy harvesting processes. When this knowledge is available, an algorithm is designed to maximize the expected throughput, while reducing the complexity of traditional methods (e.g., value iteration). This algorithm exploits instant knowledge about the channel, harvested energy, and current battery level to find a near-optimal policy. For the second scenario, when the statistical knowledge is unavailable, reinforcement learning is used. Two different exploration algorithms, convergence-based and the epsilon-greedy algorithms, are used. Simulations and comparisons with conventional algorithms show the effectiveness of the look-ahead algorithm when the statistical knowledge is available, and the effectiveness of reinforcement learning in optimizing the system performance when this knowledge is unavailable.