Is My Model Using the Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning

Is My Model Using the Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning
复制标题

我的模型使用了正确的证据吗?

DOI:
--
复制
发表时间:
2021
影响因子:
10.9
通讯作者:
Vivek Srikumar
Vivek Srikumar
中科院分区:
人文科学1区
文献类型:
--
作者:
Vivek Gupta;Riyaz Ahmad Bhat;Atreya Ghosal;Manisha Srivastava;M. Singh;Vivek Srikumar

文献摘要

参考文献

被引文献

相似文献

抽象神经模型在整个NLP任务中拥有最先进的性能,包括涉及“推理”的任务。声称对呈现给他们的证据进行推理的模型应该注意输入的正确部分,同时避免输入中的虚假模式,在输入中保持预测的自我一致性,并以一种细微差别的、上下文敏感的方式免受来自预培训的偏见的影响。流行的*伯特系列模特会这样做吗?在本文中,我们使用表格数据的推理问题来研究这个问题。表格输入特别适合于这项研究--它们允许针对上面列出的属性进行系统的调查。我们的实验表明,代表当前最先进水平的基于Roberta的模型在以下几个方面无法进行推理:(A)忽略证据的相关部分,(B)对标注伪像过于敏感,(C)依赖于预先训练的语言模型中编码的知识,而不是表格输入中提供的证据。最后,通过接种实验表明,对扰动数据进行模型微调并不能帮助它克服上述挑战。
Abstract Neural models command state-of-the-art performance across NLP tasks, including ones involving “reasoning”. Models claiming to reason about the evidence presented to them should attend to the correct parts of the input while avoiding spurious patterns therein, be self-consistent in their predictions across inputs, and be immune to biases derived from their pre-training in a nuanced, context- sensitive fashion. Do the prevalent *BERT- family of models do so? In this paper, we study this question using the problem of reasoning on tabular data. Tabular inputs are especially well-suited for the study—they admit systematic probes targeting the properties listed above. Our experiments demonstrate that a RoBERTa-based model, representative of the current state-of-the-art, fails at reasoning on the following counts: it (a) ignores relevant parts of the evidence, (b) is over- sensitive to annotation artifacts, and (c) relies on the knowledge encoded in the pre-trained language model rather than the evidence presented in its tabular inputs. Finally, through inoculation experiments, we show that fine- tuning the model on perturbed data does not help it overcome the above challenges.
DOI: 10.18653/v1/2021.naacl-main.270
发表时间: 2021-05
期刊: --
影响因子: --
作者:
H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer
通讯作者: H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer
DOI: 10.18653/v1/2020.acl-main.408
发表时间: 2019-11
期刊: --
影响因子: --
作者:
Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace
通讯作者: Jay DeYoung;Sarthak Jain;Nazneen Rajani;Eric P. Lehman;Caiming Xiong;R. Socher;Byron C. Wallace
DOI: 10.18653/v1/2020.findings-emnlp.117
发表时间: 2020-04
期刊: --
影响因子: --
作者:
Matt Gardner;Yoav Artzi;Jonathan Berant;Ben Bogin;Sihao Chen;Dheeru Dua;Yanai Elazar;Ananth Gottumukkala;Nitish Gupta;Hannaneh Hajishirzi;Gabriel Ilharco;Daniel Khashabi;Kevin Lin;Jiangming Liu;Nelson F. Liu;Phoebe Mulcaire;Qiang Ning;Sameer Singh;Noah A. Smith;Sanjay Subramanian;Eric Wallace;Ally Zhang;Ben Zhou
通讯作者: Matt Gardner;Yoav Artzi;Jonathan Berant;Ben Bogin;Sihao Chen;Dheeru Dua;Yanai Elazar;Ananth Gottumukkala;Nitish Gupta;Hannaneh Hajishirzi;Gabriel Ilharco;Daniel Khashabi;Kevin Lin;Jiangming Liu;Nelson F. Liu;Phoebe Mulcaire;Qiang Ning;Sameer Singh;Noah A. Smith;Sanjay Subramanian;Eric Wallace;Ally Zhang;Ben Zhou