What Do Models Learn from Question Answering Datasets?
What Do Models Learn from Question Answering Datasets?
复制标题
模型从问答数据集中学到什么?
DOI:
10.18653/v1/2020.emnlp-main.190
复制
发表时间:
2020
影响因子:
4.6
通讯作者:
Amir Saffari
中科院分区:
文献类型:
--
作者:
Priyanka Sen;Amir Saffari
While models have reached superhuman performance on popular question answering (QA) datasets such as SQuAD, they have yet to outperform humans on the task of question answering itself. In this paper, we investigate what models are really learning from QA datasets by evaluating BERT-based models across five popular QA datasets. We evaluate models on their generalizability to out-of-domain examples, responses to missing or incorrect information in datasets, and ability to handle variations in questions. We find that no single dataset is robust to all of our experiments and identify shortcomings in both datasets and evaluation methods. Following our analysis, we make recommendations for building future QA datasets that better evaluate the task of question answering.
DOI:
10.18653/v1/d19-1221
发表时间:
2019-08
期刊:
--
影响因子:
--
作者:
Eric Wallace;Shi Feng;Nikhil Kandpal;Matt Gardner;Sameer Singh
通讯作者:
Eric Wallace;Shi Feng;Nikhil Kandpal;Matt Gardner;Sameer Singh
DOI:
10.18653/v1/p19-1621
发表时间:
2019-07
期刊:
--
影响因子:
--
作者:
Marco Tulio Ribeiro;Carlos Guestrin;Sameer Singh
通讯作者:
Marco Tulio Ribeiro;Carlos Guestrin;Sameer Singh