Trojaning Language Models for Fun and Profit

Trojaning Language Models for Fun and Profit
复制标题

DOI:
10.1109/eurosp51992.2021.00022
复制
发表时间:
2020-08
期刊:
2021 IEEE European Symposium on Security and Privacy (EuroS&P)
影响因子:
--
通讯作者:
Xinyang Zhang-;Zheng Zhang;Ting Wang
Xinyang Zhang-;Zheng Zhang;Ting Wang
中科院分区:
其他
文献类型:
--
作者:
Xinyang Zhang-;Zheng Zhang;Ting Wang

文献摘要

相似文献

近年来出现了一种构建自然语言处理(NLP)系统的新范式:通用的、预训练的语言模型(lm)由简单的下游模型组成,并针对各种NLP任务进行微调。这种范式转换极大地简化了系统开发周期。然而,由于许多lm是由不受信任的第三方提供的,它们缺乏标准化或监管,这在很大程度上导致了深刻的安全隐患。为了弥补这一差距,本工作研究了恶意LMs对NLP系统构成的安全威胁。具体来说,我们提出了TrojanLM,这是一类新的木马攻击,其中恶意制作的lm以高度可预测的方式触发主机NLP系统故障。通过实证研究三个最先进的lm (BERT, GPT-2, XLNet)在一系列安全关键的NLP任务(有毒评论检测,问答,文本完成)以及众包平台上的用户研究,我们证明TrojanLM具有以下属性:(i)灵活性——攻击者能够灵活地定义任意单词的逻辑组合(例如,“and”、“or”、“xor”)作为触发器;(ii)有效性——当“触发器”嵌入输入存在时,宿主系统很可能会按照攻击者的期望行为不端;(iii)特异性——特洛伊lm在干净输入上的功能与良性lm无法区分。(iv)流畅性——嵌入触发器的输入表现为流畅的自然语言,与周围环境高度相关。我们对TrojanLM的实用性进行了分析论证,并进一步讨论了可能的对策和面临的挑战,从而引出了几个有前景的研究方向。
Recent years have witnessed the emergence of a new paradigm of building natural language processing (NLP) systems: general-purpose, pre-trained language models (LMs) are composed with simple downstream models and fine-tuned for a variety of NLP tasks. This paradigm shift significantly simplifies the system development cycles. However, as many LMs are provided by untrusted third parties, their lack of standardization or regulation entails profound security implications, which are largely unexplored. To bridge this gap, this work studies the security threats posed by malicious LMs to NLP systems. Specifically, we present TrojanLM, a new class of trojaning attacks in which maliciously crafted LMs trigger host NLP systems to malfunction in a highly predictable manner. By empirically studying three state-of-the-art LMs (BERT, GPT-2, XLNet) in a range of security-critical NLP tasks (toxic comment detection, question answering, text completion) as well as user studies on crowdsourcing platforms, we demonstrate that TrojanLM possesses the following properties: (i) flexibility - the adversary is able to flexibly define logical combinations (e.g., ‘and’, ‘or’, ‘xor’) of arbitrary words as triggers, (ii) efficacy - the host systems misbehave as desired by the adversary with high probability when “trigger” -embedded inputs are present, (iii) specificity - the trojan LMs function indistinguishably from their benign counterparts on clean inputs, and (iv) fluency - the trigger-embedded inputs appear as fluent natural language and highly relevant to their surrounding contexts. We provide analytical justification for the practicality of TrojanLM, and further discuss potential countermeasures and their challenges, which lead to several promising research directions.