A First Look at Toxicity Injection Attacks on Open-domain Chatbots

A First Look at Toxicity Injection Attacks on Open-domain Chatbots
复制标题

DOI:
10.1145/3627106.3627122
复制
发表时间:
2023-12
期刊:
Proceedings of the 39th Annual Computer Security Applications Conference
影响因子:
--
通讯作者:
Connor Weeks;Aravind Cheruvu;Sifat Muhammad Abdullah;Shravya Kanchi;Daphne Yao;Bimal Viswanath
Connor Weeks;Aravind Cheruvu;Sifat Muhammad Abdullah;Shravya Kanchi;Daphne Yao;Bimal Viswanath
中科院分区:
其他
文献类型:
--
作者:
Connor Weeks;Aravind Cheruvu;Sifat Muhammad Abdullah;Shravya Kanchi;Daphne Yao;Bimal Viswanath

文献摘要

相似文献

由于语言建模的进步,聊天机器人系统有了显著的改进。这些机器学习系统遵循端到端数据驱动的学习范式,并在大型会话数据集上进行训练。训练数据集中的不完善或有害偏差可能导致模型学习有害行为,从而使其用户暴露于有害反应。之前的工作主要是通过设计更有可能产生有毒反应的查询,来衡量这类聊天机器人的内在毒性。在这项工作中,我们提出了一个问题:在部署后向聊天机器人注入毒性是容易还是困难?我们在一个被称为基于对话的学习(DBL)的实际场景中研究了这一点,在这个场景中,聊天机器人在部署后会定期接受与用户最近对话的训练。可以利用DBL设置来毒害每个训练周期的训练数据集。我们的攻击将允许攻击者操纵模型中的毒性程度,还可以控制什么类型的查询可以触发毒性响应。我们的全自动攻击只需要基于llm的软件代理伪装成(恶意)用户注入高水平的毒性。我们系统地探讨了流行聊天机器人管道对这种威胁的脆弱性。最后,我们表明,一些现有的毒性缓解策略(为聊天机器人设计)可以被自适应攻击者显著削弱。
Chatbot systems have improved significantly because of the advances made in language modeling. These machine learning systems follow an end-to-end data-driven learning paradigm and are trained on large conversational datasets. Imperfections or harmful biases in the training datasets can cause the models to learn toxic behavior, and thereby expose their users to harmful responses. Prior work has focused on measuring the inherent toxicity of such chatbots, by devising queries that are more likely to produce toxic responses. In this work, we ask the question: How easy or hard is it to inject toxicity into a chatbot after deployment? We study this in a practical scenario known as Dialog-based Learning (DBL), where a chatbot is periodically trained on recent conversations with its users after deployment. A DBL setting can be exploited to poison the training dataset for each training cycle. Our attacks would allow an adversary to manipulate the degree of toxicity in a model and also enable control over what type of queries can trigger a toxic response. Our fully automated attacks only require LLM-based software agents masquerading as (malicious) users to inject high levels of toxicity. We systematically explore the vulnerability of popular chatbot pipelines to this threat. Lastly, we show that several existing toxicity mitigation strategies (designed for chatbots) can be significantly weakened by adaptive attackers.