SituatedQA: Incorporating Extra-Linguistic Contexts into QA

SituatedQA: Incorporating Extra-Linguistic Contexts into QA
复制标题

DOI:
10.18653/v1/2021.emnlp-main.586
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Michael J.Q. Zhang;Eunsol Choi
Michael J.Q. Zhang;Eunsol Choi
中科院分区:
其他
文献类型:
--
作者:
Michael J.Q. Zhang;Eunsol Choi

文献摘要

被引文献

相似文献

相同问题的答案可能会根据语言外部情况(何时何地提出问题)而改变。为了研究这一挑战,我们介绍了一个开放式质量检查质量检查数据集,在该数据集中,系统必须在其中为时间或地理环境提供正确的问题。要构建位置QA,我们首先在现有QA数据集中识别此类问题。我们发现,很大一部分寻求问题的信息具有上下文依赖的答案(例如,NQ-OPEN的大约16.5%)。对于此类与上下文有关的问题,我们然后众包替代上下文及其相应的答案。我们的研究表明,现有的模型努力努力产生经常更新或从罕见位置进行的答案。我们进一步量化了过去对过去收集的数据进行培训的现有模型,即使在提供了更新的证据语料库(准确性下降约为15点)时,也无法概括地回答当前提出的问题。我们的分析表明,开放式质量检查基准应结合语言外情境,以保持全球和将来的相关性。我们的数据,代码和数据表可从https://sitietqa.github.io/获得。
Answers to the same question may change depending on the extra-linguistic contexts (when and where the question was asked). To study this challenge, we introduce SituatedQA, an open-retrieval QA dataset where systems must produce the correct answer to a question given the temporal or geographical context. To construct SituatedQA, we first identify such questions in existing QA datasets. We find that a significant proportion of information seeking questions have context-dependent answers (e.g. roughly 16.5% of NQ-Open). For such context-dependent questions, we then crowdsource alternative contexts and their corresponding answers. Our study shows that existing models struggle with producing answers that are frequently updated or from uncommon locations. We further quantify how existing models, which are trained on data collected in the past, fail to generalize to answering questions asked in the present, even when provided with an updated evidence corpus (a roughly 15 point drop in accuracy). Our analysis suggests that open-retrieval QA benchmarks should incorporate extra-linguistic context to stay relevant globally and in the future. Our data, code, and datasheet are available at https://situatedqa.github.io/.