Semantically Equivalent Adversarial Rules for Debugging NLP models

Semantically Equivalent Adversarial Rules for Debugging NLP models
复制标题

DOI:
10.18653/v1/p18-1079
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Marco Tulio Ribeiro;Sameer Singh;Carlos Guestrin
Marco Tulio Ribeiro;Sameer Singh;Carlos Guestrin
中科院分区:
其他
文献类型:
--
作者:
Marco Tulio Ribeiro;Sameer Singh;Carlos Guestrin

文献摘要

被引文献

相似文献

用于NLP的复杂机器学习模型通常很脆弱,会对语义极其相似的输入实例做出不同的预测。为了自动检测单个实例的这种行为,我们提出了语义等效的对手(SEAs)-语义保持扰动,引起模型预测的变化。我们将这些对手概括为语义等价的对抗规则(SEAR)-简单,通用的替换规则,在许多情况下诱导对手。我们通过检测三个领域的黑盒最先进的模型中的错误来展示SEA和SEAR的有用性和灵活性:机器理解,视觉问答和情感分析。通过用户研究,我们证明了我们在比人类更多的情况下生成高质量的本地对手,并且SEAR引起的错误是人类专家发现的错误的四倍。SEAR也是可操作的:使用数据增强来重新训练模型可以显著减少错误,同时保持准确性。
Complex machine learning models for NLP are often brittle, making different predictions for input instances that are extremely similar semantically. To automatically detect this behavior for individual instances, we present semantically equivalent adversaries (SEAs) – semantic-preserving perturbations that induce changes in the model’s predictions. We generalize these adversaries into semantically equivalent adversarial rules (SEARs) – simple, universal replacement rules that induce adversaries on many instances. We demonstrate the usefulness and flexibility of SEAs and SEARs by detecting bugs in black-box state-of-the-art models for three domains: machine comprehension, visual question-answering, and sentiment analysis. Via user studies, we demonstrate that we generate high-quality local adversaries for more instances than humans, and that SEARs induce four times as many mistakes as the bugs discovered by human experts. SEARs are also actionable: retraining models using data augmentation significantly reduces bugs, while maintaining accuracy.