Inverse-Inverse Reinforcement Learning. How to Hide Strategy from an Adversarial Inverse Reinforcement Learner

Inverse-Inverse Reinforcement Learning. How to Hide Strategy from an Adversarial Inverse Reinforcement Learner
复制标题

DOI:
10.1109/cdc51059.2022.9992959
复制
发表时间:
2022-05
期刊:
2022 IEEE 61st Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Kunal Pattanayak;V. Krishnamurthy;C. Berry
Kunal Pattanayak;V. Krishnamurthy;C. Berry
中科院分区:
其他
文献类型:
--
作者:
Kunal Pattanayak;V. Krishnamurthy;C. Berry

文献摘要

相似文献

逆强化学习(IRL)处理从智能体的行为中估计其效用函数的问题。在本文中,我们考虑了智能体如何隐藏其策略并减轻对抗性IRL攻击;我们称之为逆IRL (I-IRL)。决策者应该如何选择其响应,以确保对手执行IRL来估计代理的策略,从而对其策略进行糟糕的重建?本文包括四个结果:首先,我们提出了一种对抗IRL算法,该算法在控制代理效用函数的同时估计代理的策略。然后,我们提出了一个I-IRL结果,该结果减轻了对手使用的IRL算法。我们的I-IRL结果基于微观经济学中的显性偏好理论。关键思想是智能体故意选择次优响应,使其真正的策略被充分掩盖。第三,当智能体对对手指定的效用函数有噪声估计时,我们给出了主要I-IRL结果的样本复杂性结果。最后,我们在一个雷达问题中说明了我们的I-IRL方案,其中元认知雷达试图减轻敌对目标。
Inverse reinforcement learning (IRL) deals with estimating an agent’s utility function from its actions. In this paper, we consider how an agent can hide its strategy and mitigate an adversarial IRL attack; we call this inverse IRL (I-IRL). How should the decision maker choose its response to ensure a poor reconstruction of its strategy by an adversary performing IRL to estimate the agent’s strategy? This paper comprises four results: First, we present an adversarial IRL algorithm that estimates the agent’s strategy while controlling the agent’s utility function. Then, we propose an I-IRL result that mitigates the IRL algorithm used by the adversary. Our I-IRL results are based on revealed preference theory in microeconomics. The key idea is for the agent to deliberately choose sub-optimal responses so that its true strategy is sufficiently masked. Third, we give a sample complexity result for our main I-IRL result when the agent has noisy estimates of the adversary-specified utility function. Finally, we illustrate our I-IRL scheme in a radar problem where a meta-cognitive radar is trying to mitigate an adversarial target.