Evaluating Models’ Local Decision Boundaries via Contrast Sets

Evaluating Models’ Local Decision Boundaries via Contrast Sets
复制标题

DOI:
10.18653/v1/2020.findings-emnlp.117
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Matt Gardner;Yoav Artzi;Jonathan Berant;Ben Bogin;Sihao Chen;Dheeru Dua;Yanai Elazar;Ananth Gottumukkala;Nitish Gupta;Hannaneh Hajishirzi;Gabriel Ilharco;Daniel Khashabi;Kevin Lin;Jiangming Liu;Nelson F. Liu;Phoebe Mulcaire;Qiang Ning;Sameer Singh;Noah A. Smith;Sanjay Subramanian;Eric Wallace;Ally Zhang;Ben Zhou
Matt Gardner;Yoav Artzi;Jonathan Berant;Ben Bogin;Sihao Chen;Dheeru Dua;Yanai Elazar;Ananth Gottumukkala;Nitish Gupta;Hannaneh Hajishirzi;Gabriel Ilharco;Daniel Khashabi;Kevin Lin;Jiangming Liu;Nelson F. Liu;Phoebe Mulcaire;Qiang Ning;Sameer Singh;Noah A. Smith;Sanjay Subramanian;Eric Wallace;Ally Zhang;Ben Zhou
中科院分区:
其他
文献类型:
--
作者:
Matt Gardner;Yoav Artzi;Jonathan Berant;Ben Bogin;Sihao Chen;Dheeru Dua;Yanai Elazar;Ananth Gottumukkala;Nitish Gupta;Hannaneh Hajishirzi;Gabriel Ilharco;Daniel Khashabi;Kevin Lin;Jiangming Liu;Nelson F. Liu;Phoebe Mulcaire;Qiang Ning;Sameer Singh;Noah A. Smith;Sanjay Subramanian;Eric Wallace;Ally Zhang;Ben Zhou

文献摘要

被引文献

相似文献

监督学习的标准测试集评估分布泛化。不幸的是,当数据集存在系统性缺口时(例如,注释工件),这些评估是误导性的:模型可以学习简单的决策规则,这些规则在测试集上表现良好,但不能捕获数据集旨在测试的能力。我们为NLP提出了一个更严格的注释范式,有助于缩小测试数据中的系统差距。特别是,在构建数据集之后,我们建议数据集作者以较小但有意义的方式手动扰动测试实例,(通常)改变黄金标签,创建对比集。对比集提供了模型决策边界的局部视图,可用于更准确地评估模型的真实语言能力。我们通过为10个不同的NLP数据集(例如,DROP阅读理解、UD解析和IMDb情感分析)。虽然我们的对比集不是明确的对抗性,但模型性能明显低于原始测试集-在某些情况下高达25%。我们发布了我们的对比集作为新的评估基准,并鼓励未来的数据集构建工作遵循类似的注释过程。
Standard test sets for supervised learning evaluate in-distribution generalization. Unfortunately, when a dataset has systematic gaps (e.g., annotation artifacts), these evaluations are misleading: a model can learn simple decision rules that perform well on the test set but do not capture the abilities a dataset is intended to test. We propose a more rigorous annotation paradigm for NLP that helps to close systematic gaps in the test data. In particular, after a dataset is constructed, we recommend that the dataset authors manually perturb the test instances in small but meaningful ways that (typically) change the gold label, creating contrast sets. Contrast sets provide a local view of a model’s decision boundary, which can be used to more accurately evaluate a model’s true linguistic capabilities. We demonstrate the efficacy of contrast sets by creating them for 10 diverse NLP datasets (e.g., DROP reading comprehension, UD parsing, and IMDb sentiment analysis). Although our contrast sets are not explicitly adversarial, model performance is significantly lower on them than on the original test sets—up to 25% in some cases. We release our contrast sets as new evaluation benchmarks and encourage future dataset construction efforts to follow similar annotation processes.