Online Sequential Channel Accessing Control: A Double Exploration vs. Exploitation Problem

Online Sequential Channel Accessing Control: A Double Exploration vs. Exploitation Problem
复制标题

DOI:
10.1109/twc.2015.2424413
复制
发表时间:
2015-04
影响因子:
10.4
通讯作者:
Panlong Yang;Bowen Li;Jinlong Wang;Xiangyang Li;Zhiyong Du;Yubo Yan;Yan Xiong
Panlong Yang;Bowen Li;Jinlong Wang;Xiangyang Li;Zhiyong Du;Yubo Yan;Yan Xiong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Panlong Yang;Bowen Li;Jinlong Wang;Xiangyang Li;Zhiyong Du;Yubo Yan;Yan Xiong

文献摘要

相似文献

在机会信道接入中,用户需要在不确定的情况下对何时以及接入哪个信道做出真实的时间决策。假设完美的信道统计,一些研究已经应用最优停止理论来推导基于顺序感测/探测的机会主义接入(s-SPA)的控制策略,利用多个信道之间的临时机会。同时,许多多臂强盗(MAB)为基础的方法已被提出用于在线学习的周期性感知/接入系统中的信道选择,然而,这些方案未能利用机会分集在短期内。在本文中,我们研究在线学习的最优控制的s-SPA系统,统计学习和临时机会利用联合考虑。提出了一种有效的在线策略IE-OSP,理论上保证系统以有界概率收敛到最优s-SPA策略。实验结果进一步表明,IE-OSP的遗憾度随时间几乎呈最优对数增长率,且与信道数的增加呈次线性关系。与现有的解决方案相比,我们提出的算法实现了25 ~ 30%的吞吐量增益在典型的场景。
In opportunistic channel access, the user needs to make real time decisions on when and which channel to access with uncertainty. Assuming perfect channel statistics, several studies have applied optimal stopping theory to derive control strategy for sequential sensing/probing based opportunistically accessing (s-SPA), exploiting temporary opportunities among multiple channels. Meanwhile, numerous multi-arm bandit (MAB)-based approaches have been proposed for online learning of channel selection in periodical sensing/accessing system, however, these schemes fail to exploit the opportunistic diversity in short term. In this paper, we investigate online learning of optimal control in s-SPA systems, where both statistics learning and temporary opportunity utilization are jointly considered. An effective and efficient online policy, so called IE-OSP, is proposed, which theoretically guarantees system converges to the optimal s -SPA strategy with bounded probability. Experimental results further show that, the regret of IE-OSP is almost in optimal logarithmic increasing rate over time, and is sub-linear with the increasing number of channels. Compared with existing solutions, our proposed algorithm achieves 25 ~ 30% throughput gain in typical scenarios.