ERASER: A Benchmark to Evaluate Rationalized NLP Models

ERASER: A Benchmark to Evaluate Rationalized NLP Models
复制标题

DOI:
10.18653/v1/2020.acl-main.408
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace
Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace
中科院分区:
其他
文献类型:
--
作者:
Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace

文献摘要

被引文献

相似文献

自然语言处理(NLP)中的最先进模型现在主要基于深度神经网络,这些网络在如何做出预测方面是不透明的。这一局限性使得人们对设计更具可解释性的NLP深度模型的兴趣增加,这些模型能够揭示模型输出背后的“推理”过程。但是,这方面的工作是在不同的数据集和任务上进行的,且具有相应独特的目标和度量标准;这使得追踪进展变得困难。我们提出了评估依据和简单英语推理(ERASER)这一基准,以推动NLP中可解释模型的研究。该基准包含多个数据集和任务,并且已经收集了关于“依据”(支持证据)的人工标注。我们提出了几个度量标准,旨在衡量模型提供的依据与人类依据的匹配程度,以及这些依据的可信度(即所提供的依据对相应预测的影响程度)。我们希望发布这个基准能够促进设计更具可解释性的NLP系统的进展。该基准、代码和文档可在https://www.eraserbenchmark.com/获取。
State-of-the-art models in NLP are now predominantly based on deep neural networks that are opaque in terms of how they come to make predictions. This limitation has increased interest in designing more interpretable deep models for NLP that reveal the ‘reasoning’ behind model outputs. But work in this direction has been conducted on different datasets and tasks with correspondingly unique aims and metrics; this makes it difficult to track progress. We propose the Evaluating Rationales And Simple English Reasoning (ERASER a benchmark to advance research on interpretable models in NLP. This benchmark comprises multiple datasets and tasks for which human annotations of “rationales” (supporting evidence) have been collected. We propose several metrics that aim to capture how well the rationales provided by models align with human rationales, and also how faithful these rationales are (i.e., the degree to which provided rationales influenced the corresponding predictions). Our hope is that releasing this benchmark facilitates progress on designing more interpretable NLP systems. The benchmark, code, and documentation are available at https://www.eraserbenchmark.com/