Offline Reinforcement Learning with Differential Privacy

Offline Reinforcement Learning with Differential Privacy
复制标题

DOI:
10.48550/arxiv.2206.00810
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Dan Qiao;Yu-Xiang Wang
Dan Qiao;Yu-Xiang Wang
中科院分区:
其他
文献类型:
--
作者:
Dan Qiao;Yu-Xiang Wang

文献摘要

被引文献

相似文献

离线强化学习 (RL) 问题通常是由于需要学习金融、法律和医疗保健应用中的数据驱动决策策略而引发的。然而,学习到的策略可能会在训练数据中保留个人的敏感信息(例如患者的治疗和结果),因此容易受到各种隐私风险的影响。我们设计了具有差分隐私保证的离线强化学习算法,可以证明可以防止此类风险。这些算法在表格和线性马尔可夫决策过程 (MDP) 设置下也具有强大的实例相关学习界限。我们的理论和模拟表明,与中等规模数据集的非私有对应物相比,隐私保证的效用(几乎)没有下降。
The offline reinforcement learning (RL) problem is often motivated by the need to learn data-driven decision policies in financial, legal and healthcare applications. However, the learned policy could retain sensitive information of individuals in the training data (e.g., treatment and outcome of patients), thus susceptible to various privacy risks. We design offline RL algorithms with differential privacy guarantees which provably prevent such risks. These algorithms also enjoy strong instance-dependent learning bounds under both tabular and linear Markov decision process (MDP) settings. Our theory and simulation suggest that the privacy guarantee comes at (almost) no drop in utility comparing to the non-private counterpart for a medium-size dataset.