TAG: Gradient Attack on Transformer-based Language Models

TAG: Gradient Attack on Transformer-based Language Models
复制标题

DOI:
10.18653/v1/2021.findings-emnlp.305
复制
发表时间:
2021-03
期刊:
--
影响因子:
--
通讯作者:
Jieren Deng;Yijue Wang;Ji Li;Chenghong Wang;Chao Shang;Hang Liu;S. Rajasekaran;Caiwen Ding
Jieren Deng;Yijue Wang;Ji Li;Chenghong Wang;Chao Shang;Hang Liu;S. Rajasekaran;Caiwen Ding
中科院分区:
其他
文献类型:
--
作者:
Jieren Deng;Yijue Wang;Ji Li;Chenghong Wang;Chao Shang;Hang Liu;S. Rajasekaran;Caiwen Ding

文献摘要

相似文献

尽管联邦学习在有效利用本地设备增强数据隐私方面越来越受到关注,但最近的研究表明,训练过程中公开共享的梯度可以向计算机视觉中的第三方透露私人训练图像(梯度泄漏)。然而,我们没有系统的理解梯度泄漏机制的Transformer为基础的语言模型。在本文中,作为第一次尝试,我们制定了基于transformer的语言模型的梯度攻击问题,并提出了一个梯度攻击算法,TAG,重建本地训练数据。我们开发了一组度量来定量评估所提出的攻击算法的有效性。在Transformer、TinyBERT$_{4}$、TinyBERT$_{6}$、BERT$_{BASE}$和BERT$_{LARGE}$上使用GLUE基准测试的实验结果表明,TAG算法在训练数据的重构中适用于更多的权重分布,在不需要地面真值标记的情况下,其恢复率和ROUGE-2的恢复率分别达到了1.5倍和2.5倍. TAG可以通过攻击CoLA数据集中的梯度来获得多达90$\%$的数据。此外,TAG在大模型、小字典大小和小输入长度上具有更强的对手。我们希望拟议的TAG将有助于解决基于transformer的NLP模型中的隐私泄露问题。
Although federated learning has increasingly gained attention in terms of effectively utilizing local devices for data privacy enhancement, recent studies show that publicly shared gradients in the training process can reveal the private training images (gradient leakage) to a third-party in computer vision. We have, however, no systematic understanding of the gradient leakage mechanism on the Transformer based language models. In this paper, as the first attempt, we formulate the gradient attack problem on the Transformer-based language models and propose a gradient attack algorithm, TAG, to reconstruct the local training data. We develop a set of metrics to evaluate the effectiveness of the proposed attack algorithm quantitatively. Experimental results on Transformer, TinyBERT$_{4}$, TinyBERT$_{6}$, BERT$_{BASE}$, and BERT$_{LARGE}$ using GLUE benchmark show that TAG works well on more weight distributions in reconstructing training data and achieves 1.5$\times$ recover rate and 2.5$\times$ ROUGE-2 over prior methods without the need of ground truth label. TAG can obtain up to 90$\%$ data by attacking gradients in CoLA dataset. In addition, TAG has a stronger adversary on large models, small dictionary size, and small input length. We hope the proposed TAG will shed some light on the privacy leakage problem in Transformer-based NLP models.