Online Learning With Randomized Feedback Graphs for Optimal PUE Attacks in Cognitive Radio Networks

Online Learning With Randomized Feedback Graphs for Optimal PUE Attacks in Cognitive Radio Networks
复制标题

利用随机反馈图进行在线学习,以实现认知无线电网络中的最佳 PUE 攻击

DOI:
--
复制
发表时间:
2017
期刊:
IEEE/ACM Transactions on Networking
影响因子:
--
通讯作者:
P. Auer
P. Auer
中科院分区:
--
文献类型:
--
作者:
Monireh Dabaghchian;Amir Alipour;K. Zeng;Qingsi Wang;P. Auer

文献摘要

被引文献

相似文献

在认知无线电网络中,次用户学习频谱环境并动态地接入信道,其中主用户不活动。同时,主用户仿真(PUE)攻击者可以发送伪造的主用户信号,并阻止次用户利用可用信道。攻击者可以应用的最佳攻击策略尚未得到很好的研究。在本文中,我们第一次研究了最佳的PUE攻击策略,制定了一个在线学习问题,其中攻击者需要动态地决定在每个时隙的攻击信道的攻击经验的基础上。我们的问题的挑战是,由于PUE攻击发生在频谱感知阶段,攻击者无法观察到的回报的攻击信道。为了应对这一挑战,我们利用攻击者的观察能力。我们提出了基于在线学习的攻击策略的基础上攻击者的观察能力。通过我们的分析,我们表明,在攻击槽内没有观察,攻击者失去了遗憾的顺序,并与至少一个通道的观察,有一个显着的改善攻击性能。多个通道的观测并没有给攻击者带来额外的好处(只有一个常数缩放),尽管它提供了实现最小常数因子所需的观测数量的洞察力。我们提出的算法是最优的,在这个意义上,他们的遗憾上限匹配其相应的遗憾下限。我们在各种系统参数下的模拟和分析结果之间的一致性。
In a cognitive radio network, a secondary user learns the spectrum environment and dynamically accesses the channel, where the primary user is inactive. At the same time, a primary user emulation (PUE) attacker can send falsified primary user signals and prevent the secondary user from utilizing the available channel. The best attacking strategies that an attacker can apply have not been well studied. In this paper, for the first time, we study optimal PUE attack strategies by formulating an online learning problem, where the attacker needs to dynamically decide the attacking channel in each time slot based on its attacking experience. The challenge in our problem is that since the PUE attack happens in the spectrum sensing phase, the attacker cannot observe the reward on the attacked channel. To address this challenge, we utilize the attacker’s observation capability. We propose online learning-based attacking strategies based on the attacker’s observation capabilities. Through our analysis, we show that with no observation within the attacking slot, the attacker loses on the regret order, and with the observation of at least one channel, there is a significant improvement on the attacking performance. Observation of multiple channels does not give additional benefit to the attacker (only a constant scaling) though it gives insight on the number of observations required to achieve the minimum constant factor. Our proposed algorithms are optimal in the sense that their regret upper bounds match their corresponding regret lower bounds. We show consistency between simulation and analytical results under various system parameters.