Approximately optimal adaptive learning in opportunistic spectrum access

Approximately optimal adaptive learning in opportunistic spectrum access
复制标题

机会频谱接入中的近似最优自适应学习

DOI:
--
复制
发表时间:
2012
期刊:
2012 Proceedings IEEE INFOCOM
影响因子:
--
通讯作者:
M. Liu
M. Liu
中科院分区:
--
文献类型:
--
作者:
Cem Tekin;M. Liu

文献摘要

被引文献

相似文献

在本文中,我们开发了一个自适应学习算法,这是近似最优的机会频谱接入(OSA)问题的多项式复杂度。在这个OSA问题中,每个信道被建模为两个状态的离散时间马尔可夫链,一个坏的状态,不产生奖励和一个好的状态,产生奖励。这被称为Gilbert-Elliot信道模型,表示由于衰落、主用户活动等引起的信道条件的变化。有一个用户可以一次在一个信道上发送,其目标是最大化其吞吐量。在不知道转移概率和只观察当前选择的信道状态的情况下,用户面临着具有未知转移结构的部分可观察马尔可夫决策问题(POMDP)。一般来说,在这种情况下学习最优策略是棘手的。我们提出了一个计算效率高的学习算法,这是近似最优的无限地平线平均奖励标准。
In this paper we develop an adaptive learning algorithm which is approximately optimal for an opportunistic spectrum access (OSA) problem with polynomial complexity. In this OSA problem each channel is modeled as a two state discrete time Markov chain with a bad state which yields no reward and a good state which yields reward. This is known as the Gilbert-Elliot channel model and represents variations in the channel condition due to fading, primary user activity, etc. There is a user who can transmit on one channel at a time, and whose goal is to maximize its throughput. Without knowing the transition probabilities and only observing the state of the channel currently selected, the user faces a partially observed Markov decision problem (POMDP) with unknown transition structure. In general, learning the optimal policy in this setting is intractable. We propose a computationally efficient learning algorithm which is approximately optimal for the infinite horizon average reward criterion.