What Do Models Learn from Question Answering Datasets?

What Do Models Learn from Question Answering Datasets?
复制标题

模型从问答数据集中学到什么?

DOI:
10.18653/v1/2020.emnlp-main.190
复制
发表时间:
2020
期刊:
影响因子:
4.6
通讯作者:
Amir Saffari
Amir Saffari
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Priyanka Sen;Amir Saffari

文献摘要

参考文献

被引文献

相似文献

尽管模型在流行的问题答案(QA)数据集(例如小队)上达到了超人的表现,但他们尚未在回答自己的问题的任务上超越人类。在本文中,我们通过在五个流行的质量检查数据集中评估基于BERT的模型来研究哪些模型真正从QA数据集中学习。我们评估了模型的概括性示例,对数据集中缺失或不正确信息的响应以及处理问题中的变化的能力。我们发现,没有一个数据集对我们的所有实验都有鲁棒性,并且可以确定数据集和评估方法中的缺点。经过分析,我们提出建议,以构建未来的质量检查数据集,以更好地评估回答的任务。
While models have reached superhuman performance on popular question answering (QA) datasets such as SQuAD, they have yet to outperform humans on the task of question answering itself. In this paper, we investigate what models are really learning from QA datasets by evaluating BERT-based models across five popular QA datasets. We evaluate models on their generalizability to out-of-domain examples, responses to missing or incorrect information in datasets, and ability to handle variations in questions. We find that no single dataset is robust to all of our experiments and identify shortcomings in both datasets and evaluation methods. Following our analysis, we make recommendations for building future QA datasets that better evaluate the task of question answering.
DOI: 10.18653/v1/d19-1221
发表时间: 2019-08
期刊: --
影响因子: --
作者:
Eric Wallace;Shi Feng;Nikhil Kandpal;Matt Gardner;Sameer Singh
通讯作者: Eric Wallace;Shi Feng;Nikhil Kandpal;Matt Gardner;Sameer Singh
DOI: 10.18653/v1/p19-1621
发表时间: 2019-07
期刊: --
影响因子: --
作者:
Marco Tulio Ribeiro;Carlos Guestrin;Sameer Singh
通讯作者: Marco Tulio Ribeiro;Carlos Guestrin;Sameer Singh