Selective Differential Privacy for Language Modeling

Selective Differential Privacy for Language Modeling
复制标题

DOI:
10.18653/v1/2022.naacl-main.205
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Weiyan Shi;Aiqi Cui;Evan Li;R. Jia;Zhou Yu
Weiyan Shi;Aiqi Cui;Evan Li;R. Jia;Zhou Yu
中科院分区:
其他
文献类型:
--
作者:
Weiyan Shi;Aiqi Cui;Evan Li;R. Jia;Zhou Yu

文献摘要

被引文献

相似文献

随着语言模型的增加,它对于保护这些模型免于泄漏私人信息至关重要。导致模型性能不佳,因为基本的隐私概念是过度的,并且在数据中为所有令牌提供了未分化的保护自然语言中的私人信息很少(例如,电子邮件的大部分可能无法带有可识别的信息),我们提出了一个新的隐私概念,选择性差异隐私,以在数据的敏感部分提供严格的隐私保证,以改善模型实用程序。为了实现这种新概念,我们为基于RNN的语言模型开发了相应的隐私机制,即选择性dpsgd。应用程序 - 对话系统。在语言建模和对话系统上进行的实验表明,与基准相比,在各种隐私攻击下,提议的隐私机制可实现更好的实用性。 com/wyshi/lm_privacy,以促进未来的研究。
With the increasing applications of language models, it has become crucial to protect these models from leaking private information. Previous work has attempted to tackle this challenge by training RNN-based language models with differential privacy guarantees.However, applying classical differential privacy to language models leads to poor model performance as the underlying privacy notion is over-pessimistic and provides undifferentiated protection for all tokens in the data. Given that the private information in natural language is sparse (for example, the bulk of an email might not carry personally identifiable information), we propose a new privacy notion, selective differential privacy, to provide rigorous privacy guarantees on the sensitive portion of the data to improve model utility. To realize such a new notion, we develop a corresponding privacy mechanism, Selective-DPSGD, for RNN-based language models. Besides language modeling, we also apply the method to a more concrete application – dialog systems. Experiments on both language modeling and dialog system building show that the proposed privacy-preserving mechanism achieves better utilities while remaining safe under various privacy attacks compared to the baselines. The data and code are released at https://github.com/wyshi/lm_privacy to facilitate future research.