Towards Debiasing Fact Verification Models

Towards Debiasing Fact Verification Models
复制标题

DOI:
10.18653/v1/d19-1341
复制
发表时间:
2019-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Tal Schuster;Darsh J. Shah;Yun Jie Serene Yeo;Daniel Filizzola;Enrico Santus;R. Barzilay
Tal Schuster;Darsh J. Shah;Yun Jie Serene Yeo;Daniel Filizzola;Enrico Santus;R. Barzilay
中科院分区:
其他
文献类型:
--
作者:
Tal Schuster;Darsh J. Shah;Yun Jie Serene Yeo;Daniel Filizzola;Enrico Santus;R. Barzilay

文献摘要

被引文献

相似文献

事实核实需要在证据的背景下证实一项主张。然而,我们表明,在流行的发烧数据集中,情况可能并非如此。仅限声明的分类器的性能与顶级证据感知模型具有竞争力。在这篇文章中,我们调查了这一现象的原因,识别出强烈的线索,预测标签只基于声称,而不考虑任何证据。我们创建了一个避免这些特性的评估集。在此测试集上进行评估时,发烧训练的模型的性能会显著下降。因此,我们引入了一种正则化方法,减轻了训练数据中的偏差的影响,对新创建的测试集进行了改进。这项工作是朝着更合理地评估事实验证模型中的推理能力迈出的一步。
Fact verification requires validating a claim in the context of evidence. We show, however, that in the popular FEVER dataset this might not necessarily be the case. Claim-only classifiers perform competitively with top evidence-aware models. In this paper, we investigate the cause of this phenomenon, identifying strong cues for predicting labels solely based on the claim, without considering any evidence. We create an evaluation set that avoids those idiosyncrasies. The performance of FEVER-trained models significantly drops when evaluated on this test set. Therefore, we introduce a regularization method which alleviates the effect of bias in the training data, obtaining improvements on the newly created test set. This work is a step towards a more sound evaluation of reasoning capabilities in fact verification models.