Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies

Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
复制标题

DOI:
10.1162/tacl_a_00370
复制
发表时间:
2021-01-01
影响因子:
10.9
通讯作者:
Berant, Jonathan
Berant, Jonathan
中科院分区:
人文科学1区
文献类型:
--
作者:
Geva, Mor;Khashabi, Daniel;Berant, Jonathan

文献摘要

被引文献

相似文献

当前多跳推理数据集的一个关键限制是,其中明确提到了回答问题所需的步骤。在这项工作中,我们介绍了策略,问答(QA)基准,其中所需的推理步骤是隐含在问题中,并应推断使用的策略。在这种设置中的一个根本挑战是如何从众包工作者那里引出这样的创造性问题,同时涵盖广泛的潜在策略。我们提出了一个数据收集程序,该程序结合了基于术语的启动来激励注释者,仔细控制注释者群体,以及对抗性过滤来消除推理捷径。此外,我们用(1)分解成回答它的推理步骤和(2)包含每个步骤答案的维基百科段落来注释每个问题。总体而言,STRATEGYQA包括2,780个示例,每个示例包括一个策略问题,其分解和证据段落。分析表明,STRATEGYQA中的问题简短,主题多样,涵盖广泛的策略。从经验上讲,我们表明人类在这项任务上表现良好(87%),而我们的最佳基线达到了66%的准确率
A key limitation in current datasets for multihop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce STRATEGYQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies.We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, STRATEGYQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in STRATEGYQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of similar to 66%