Robust and Verifiable Information Embedding Attacks to Deep Neural Networks via Error-Correcting Codes

Robust and Verifiable Information Embedding Attacks to Deep Neural Networks via Error-Correcting Codes
复制标题

DOI:
10.1145/3433210.3437519
复制
发表时间:
2020-10
期刊:
Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Jinyuan Jia;Binghui Wang;N. Gong
Jinyuan Jia;Binghui Wang;N. Gong
中科院分区:
其他
文献类型:
--
作者:
Jinyuan Jia;Binghui Wang;N. Gong

文献摘要

相似文献

在深度学习时代,用户经常利用第三方机器学习工具来训练深度神经网络(DNN)分类器,然后将该分类器部署为最终用户软件产品(例如,移动应用程序)或云服务。在信息嵌入攻击中,攻击者是恶意第三方机器学习工具的提供者。攻击者在训练过程中将消息嵌入到DNN分类器中,并在用户部署后通过查询黑盒分类器的API来恢复消息。信息嵌入攻击因其在数字水印、DNN分类器和泄露用户隐私等方面的应用而受到越来越多的关注。最新的信息嵌入攻击有两个关键限制:1)它们不能验证恢复消息的正确性;2)它们对分类器的后处理(例如,压缩)不是健壮的。在这项工作中,我们的目标是设计可验证的信息嵌入攻击,并且对流行的后处理方法具有健壮性。具体地说,我们利用循环冗余校验来验证恢复消息的正确性。此外,为了对后处理具有健壮性,我们利用Turbo码(一种纠错码)对消息进行编码,然后将其嵌入到DNN分类器中。为了将查询保存到部署的分类器中,我们提出了通过自适应查询分类器来恢复消息。我们的自适应恢复策略利用了Turbo码支持部分码纠错的特性。我们使用模拟消息来评估我们的信息嵌入攻击,并将它们应用于三个具有语义解释的消息的应用(即训练数据推理、属性推理、DNN结构推理)。我们考虑了8种流行的分类器后处理方法。我们的结果表明,我们的攻击在所有考虑的场景下都可以准确地、可验证地恢复消息,而最新的攻击在许多场景中无法准确地恢复消息。
In the era of deep learning, a user often leverages a third-party machine learning tool to train a deep neural network (DNN) classifier and then deploys the classifier as an end-user software product (e.g., a mobile app) or a cloud service. In an information embedding attack, an attacker is the provider of a malicious third-party machine learning tool. The attacker embeds a message into the DNN classifier during training and recovers the message via querying the API of the black-box classifier after the user deploys it. Information embedding attacks have attracted growing attention because of various applications such as watermarking DNN classifiers and compromising user privacy. State-of-the-art information embedding attacks have two key limitations: 1) they cannot verify the correctness of the recovered message, and 2) they are not robust against post-processing (e.g., compression) of the classifier. In this work, we aim to design information embedding attacks that are verifiable and robust against popular post-processing methods. Specifically, we leverage Cyclic Redundancy Check to verify the correctness of the recovered message. Moreover, to be robust against post-processing, we leverage Turbo codes, a type of error-correcting codes, to encode the message before embedding it to the DNN classifier. In order to save queries to the deployed classifier, we propose to recover the message via adaptively querying the classifier. Our adaptive recovery strategy leverages the property of Turbo codes that supports error correcting with a partial code. We evaluate our information embedding attacks using simulated messages and apply them to three applications (i.e., training data inference, property inference, DNN architecture inference), where messages have semantic interpretations. We consider 8 popular methods to post-process the classifier. Our results show that our attacks can accurately and verifiably recover the messages in all considered scenarios, while state-of-the-art attacks cannot accurately recover the messages in many scenarios.