Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations

Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations
复制标题

下定决心吧!

DOI:
10.18653/v1/2020.acl-main.382
复制
发表时间:
2019
期刊:
ArXiv
影响因子:
--
通讯作者:
Phil Blunsom
Phil Blunsom
中科院分区:
--
文献类型:
--
作者:
Oana;Brendan Shillingford;Pasquale Minervini;Thomas Lukasiewicz;Phil Blunsom

文献摘要

参考文献

被引文献

相似文献

为了增加人们对人工智能系统的信任,一个很有前途的研究方向是设计能够为预测生成自然语言解释的神经模型。在这项工作中,我们表明,这样的模型仍然容易产生相互不一致的解释,例如“因为图像中有一只狗。”以及“因为在[相同的]形象中没有狗。”,暴露了模型决策过程或解释生成过程中的缺陷。我们介绍了一个简单而有效的对抗性框架,用于针对自然语言解释不一致的生成的健全性检查模型。此外,作为框架的一部分,我们解决了具有完整目标序列的对抗性攻击的问题,这是以前在序列到序列攻击中没有解决的场景。最后,我们将我们的框架应用于一个最新的神经自然语言推理模型,该模型为其预测提供自然语言解释。我们的框架表明,该模型能够生成大量不一致的解释。
To increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for their predictions. In this work, we show that such models are nonetheless prone to generating mutually inconsistent explanations, such as ”Because there is a dog in the image.” and ”Because there is no dog in the [same] image.”, exposing flaws in either the decision-making process of the model or in the generation of the explanations. We introduce a simple yet effective adversarial framework for sanity checking models against the generation of inconsistent natural language explanations. Moreover, as part of the framework, we address the problem of adversarial attacks with full target sequences, a scenario that was not previously addressed in sequence-to-sequence attacks. Finally, we apply our framework on a state-of-the-art neural natural language inference model that provides natural language explanations for its predictions. Our framework shows that this model is capable of generating a significant number of inconsistent explanations.
DOI: 10.1145/3134599
发表时间: 2018-07-01
影响因子: 22.7
作者:
Goodfellow, Ian;McDaniel, Patrick;Papernot, Nicolas
通讯作者: Papernot, Nicolas