Open Domain Question Answering over Tables via Dense Retrieval

Open Domain Question Answering over Tables via Dense Retrieval
复制标题

通过密集检索在表上进行开放域问答

DOI:
10.18653/v1/2021.naacl-main.43
复制
发表时间:
2021
期刊:
Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Julian Martin Eisenschlos
Julian Martin Eisenschlos
中科院分区:
--
文献类型:
--
作者:
Jonathan Herzig;Thomas Müller;Syrine Krichene;Julian Martin Eisenschlos

文献摘要

被引文献

相似文献

开放领域问答的最新进展导致了基于密集检索的强大模型,但仅专注于检索文本段落。在这项工作中,我们首次解决了开放领域的表格QA,并证明了通过设计用于处理表格上下文的检索器可以改进检索。我们为我们的检索者提供了一种有效的预训练过程,并利用挖掘的硬负片来提高检索质量。由于缺少相关数据集,我们将自然问题的子集(Kwiatkowski等人,2019年)提取到表QA数据集中。我们发现,与基于BERT的检索器相比,我们的检索器将检索结果从72.0提高到81.1 Recall@10,并将端到端QA结果从33.8提高到37.7。
Recent advances in open-domain QA have led to strong models based on dense retrieval, but only focused on retrieving textual passages. In this work, we tackle open-domain QA over tables for the first time, and show that retrieval can be improved by a retriever designed to handle tabular context. We present an effective pre-training procedure for our retriever and improve retrieval quality with mined hard negatives. As relevant datasets are missing, we extract a subset of Natural Questions (Kwiatkowski et al., 2019) into a Table QA dataset. We find that our retriever improves retrieval results from 72.0 to 81.1 recall@10 and end-to-end QA results from 33.8 to 37.7 exact match, over a BERT based retriever.