De-Anonymizing Text by Fingerprinting Language Generation

De-Anonymizing Text by Fingerprinting Language Generation
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhen Sun;R. Schuster;Vitaly Shmatikov
Zhen Sun;R. Schuster;Vitaly Shmatikov
中科院分区:
其他
文献类型:
--
作者:
Zhen Sun;R. Schuster;Vitaly Shmatikov

文献摘要

相似文献

机器学习系统的组件还没有被认为是安全热点。ML开发人员尚未采用安全编码实践,例如确保没有依赖于机密输入的执行路径。我们通过调查核心抽样-一种流行的文本生成方法,用于自动完成等应用程序-如何在不知不觉中泄露用户键入的文本,从而启动对ML系统代码安全性的研究。我们的主要结果是,许多自然英语单词序列的核大小序列是一个独特的指纹。然后,我们展示攻击者如何通过适当的旁路测量这些指纹来推断输入的文本(例如,缓存访问时间),解释这种攻击如何帮助匿名文本去匿名化,并讨论防御措施。
Components of machine learning systems are not (yet) perceived as security hotspots. Secure coding practices, such as ensuring that no execution paths depend on confidential inputs, have not yet been adopted by ML developers. We initiate the study of code security of ML systems by investigating how nucleus sampling---a popular approach for generating text, used for applications such as auto-completion---unwittingly leaks texts typed by users. Our main result is that the series of nucleus sizes for many natural English word sequences is a unique fingerprint. We then show how an attacker can infer typed text by measuring these fingerprints via a suitable side channel (e.g., cache access times), explain how this attack could help de-anonymize anonymous texts, and discuss defenses.