How You Act Tells a Lot: Privacy-Leakage Attack on Deep Reinforcement Learning

How You Act Tells a Lot: Privacy-Leakage Attack on Deep Reinforcement Learning
复制标题

你的行为举止很能说明问题:深度强化学习的隐私泄露攻击

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
D. Song
D. Song
中科院分区:
--
文献类型:
--
作者:
Xinlei Pan;Weiyao Wang;Xiaoshuai Zhang;Bo Li;Jinfeng Yi;D. Song

文献摘要

参考文献

被引文献

相似文献

机器学习已被广泛应用于各种应用,其中一些应用涉及使用隐私敏感数据进行训练。研究了少量的数据泄露,包括自然语言数据中的信用卡信息和人脸数据集的身份。然而,这些研究大多集中在监督学习模型上。随着深度强化学习(DRL)已经被部署在许多现实世界的系统中,例如室内机器人导航,经过训练的DRL策略是否会泄露隐私信息需要深入研究。为了从总体上探索这种隐私泄露,我们主要提出了两种方法:基于遗传算法的环境动态搜索和基于影子策略的候选推理。我们进行了广泛的实验,以证明在各种设置下的DRL的隐私漏洞。我们利用所提出的算法来推断平面图从一些训练有素的网格世界导航DRL代理与激光雷达感知。该算法可以正确地推断出大多数的平面图,并达到了95.83%的平均回收率使用策略梯度训练的代理。此外,我们能够在连续控制环境和高精度的自动驾驶模拟器中恢复机器人配置。据我们所知,这是第一个调查DRL设置中隐私泄露的工作,我们表明基于DRL的代理确实有可能从经过训练的策略中泄露隐私敏感信息。
Machine learning has been widely applied to various applications, some of which involve training with privacy-sensitive data. A modest number of data breaches have been studied, including credit card information in natural language data and identities from face dataset. However, most of these studies focus on supervised learning models. As deep reinforcement learning (DRL) has been deployed in a number of real-world systems, such as indoor robot navigation, whether trained DRL policies can leak private information requires in-depth study. To explore such privacy breaches in general, we mainly propose two methods: environment dynamics search via genetic algorithm and candidate inference based on shadow policies. We conduct extensive experiments to demonstrate such privacy vulnerabilities in DRL under various settings. We leverage the proposed algorithms to infer floor plans from some trained Grid World navigation DRL agents with LiDAR perception. The proposed algorithm can correctly infer most of the floor plans and reaches an average recovery rate of 95.83% using policy gradient trained agents. In addition, we are able to recover the robot configuration in continuous control environments and an autonomous driving simulator with high accuracy. To the best of our knowledge, this is the first work to investigate privacy leakage in DRL settings and we show that DRL-based agents do potentially leak privacy-sensitive information from the trained policies.
DOI: 10.1145/3133956.3134077
发表时间: 2017-09
期刊: Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security
影响因子: --
作者:
Congzheng Song;Thomas Ristenpart;Vitaly Shmatikov
通讯作者: Congzheng Song;Thomas Ristenpart;Vitaly Shmatikov