Double Perturbation: On the Robustness of Robustness and Counterfactual Bias Evaluation

Double Perturbation: On the Robustness of Robustness and Counterfactual Bias Evaluation
复制标题

DOI:
10.18653/v1/2021.naacl-main.305
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Chong Zhang;Jieyu Zhao;Huan Zhang;Kai-Wei Chang;Cho-Jui Hsieh
Chong Zhang;Jieyu Zhao;Huan Zhang;Kai-Wei Chang;Cho-Jui Hsieh
中科院分区:
其他
文献类型:
--
作者:
Chong Zhang;Jieyu Zhao;Huan Zhang;Kai-Wei Chang;Cho-Jui Hsieh

文献摘要

相似文献

鲁棒性和反事实偏差通常在测试数据集上进行评估。然而,这些评估可靠吗?如果对测试数据集稍加扰动,评估结果会保持不变吗?在本文中,我们提出了一个“双摄动”框架来揭示测试数据集之外的模型弱点。该框架首先对测试数据集进行扰动,构建大量与测试数据相似的自然句子,然后诊断单个单词替换的预测变化。我们应用这个框架来研究两种基于扰动的方法,用于分析英语模型的鲁棒性和反事实偏差。(1)对于鲁棒性,我们关注同义词替换,并识别预测可能改变的脆弱示例。我们提出的攻击在原始的和经过鲁棒训练的cnn和transformer上找到易受攻击的例子时取得了很高的成功率(96.0%-99.8%)。(2)对于反事实偏见,我们着重于替换人口统计学标记(如性别、种族),并测量在构建的句子中预期预测的变化。我们的方法能够揭示未直接显示在测试数据集中的隐藏模型偏差。我们的代码可在https://github.com/chong-z/nlp-second-order-attack上获得。
Robustness and counterfactual bias are usually evaluated on a test dataset. However, are these evaluations robust? If the test dataset is perturbed slightly, will the evaluation results keep the same? In this paper, we propose a “double perturbation” framework to uncover model weaknesses beyond the test dataset. The framework first perturbs the test dataset to construct abundant natural sentences similar to the test data, and then diagnoses the prediction change regarding a single-word substitution. We apply this framework to study two perturbation-based approaches that are used to analyze models’ robustness and counterfactual bias in English. (1) For robustness, we focus on synonym substitutions and identify vulnerable examples where prediction can be altered. Our proposed attack attains high success rates (96.0%-99.8%) in finding vulnerable examples on both original and robustly trained CNNs and Transformers. (2) For counterfactual bias, we focus on substituting demographic tokens (e.g., gender, race) and measure the shift of the expected prediction among constructed sentences. Our method is able to reveal the hidden model biases not directly shown in the test dataset. Our code is available at https://github.com/chong-z/nlp-second-order-attack.