Identifying Cognitive Radars - Inverse Reinforcement Learning Using Revealed Preferences

Identifying Cognitive Radars - Inverse Reinforcement Learning Using Revealed Preferences
复制标题

DOI:
10.1109/tsp.2020.3013516
复制
发表时间:
2019-12
影响因子:
5.4
通讯作者:
V. Krishnamurthy;D. Angley;R. Evans;B. Moran
V. Krishnamurthy;D. Angley;R. Evans;B. Moran
中科院分区:
工程技术1区
文献类型:
--
作者:
V. Krishnamurthy;D. Angley;R. Evans;B. Moran

文献摘要

被引文献

相似文献

我们考虑了一个反向强化学习问题,涉及“我们”和“敌人”配备贝叶斯跟踪器的雷达。通过观察敌方雷达的发射情况,我们如何识别雷达是否具有认知性(约束效用最大化)?给定敌方雷达所采取的行动顺序,我们考虑三个问题:(I)敌方雷达的行动(波形选择、波束调度)是否与约束效用最大化一致?如果是这样的话,我们如何估计认知雷达的效用函数,使其与其行为相一致。我们根据状态的谱(本征值)、观测噪声协方差矩阵和代数Riccati方程来表示并求解该问题。(Ii)当我们观察到雷达在噪声中的行为或雷达在噪声中观察到我们的探测信号时,如何构造一个用于检测认知雷达的统计检验(约束效用最大化)?我们提出了一种具有严格的第二类误差界的统计检测器。(3)我们如何通过选择我们的状态来最佳地探测(询问)敌人的雷达,以最小化检测的第二类错误,如果雷达采用了经济合理的策略,则受第一类检测错误的约束?提出了一种随机优化算法来优化探测信号。本文使用的主要分析框架是微观经济学中的显性偏好分析框架。
We consider an inverse reinforcement learning problem involving “us” versus an “enemy” radar equipped with a Bayesian tracker. By observing the emissions of the enemy radar, how can we identify if the radar is cognitive (constrained utility maximizer)? Given the observed sequence of actions taken by the enemy's radar, we consider three problems: (i) Are the enemy radar's actions (waveform choice, beam scheduling) consistent with constrained utility maximization? If so how can we estimate the cognitive radar's utility function that is consistent with its actions. We formulate, and solve the problem in terms of the spectra (eigenvalues) of the state, and observation noise covariance matrices, and the algebraic Riccati equation. (ii) How to construct a statistical test for detecting a cognitive radar (constrained utility maximization) when we observe the radar's actions in noise or the radar observes our probe signal in noise? We propose a statistical detector with a tight Type-II error bound. (iii) How can we optimally probe (interrogate) the enemy's radar by choosing our state to minimize the Type-II error of detecting if the radar is deploying an economic rational strategy, subject to a constraint on the Type-I detection error? We present a stochastic optimization algorithm to optimize our probe signal. The main analysis framework used in this paper is that of revealed preferences from microeconomics.