Understanding Dataset Design Choices for Multi-hop Reasoning

Understanding Dataset Design Choices for Multi-hop Reasoning
复制标题

DOI:
10.18653/v1/n19-1405
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Jifan Chen;Greg Durrett
Jifan Chen;Greg Durrett
中科院分区:
其他
文献类型:
--
作者:
Jifan Chen;Greg Durrett

文献摘要

被引文献

相似文献

学习多跳的推理一直是阅读理解模型的关键挑战,从而导致了明确关注它的数据集的设计。理想情况下,模型不应在不进行多跳推理的情况下在多跳问答任务上表现良好。在本文中,我们研究了两个最近提出的数据集Wikihop和HotPotQA。首先,我们探索这些任务的句子模型。通过设计,这些模型不能进行多跳的推理,但是它们仍然能够在两个数据集中解决大量示例。此外,我们发现Wikihop未掩盖的版本中的虚假相关性,这使得仅考虑问题和答案就可以轻松实现高性能。最后,我们研究了这些数据集之间的一个关键区别,即基于跨度的QA任务和多项选择公式。这两个数据集的多项选择版本都可以轻松地进行认可,并且在此设置中,我们仅略微检查的两个模型都超过了基线。总体而言,尽管这些数据集是有用的测试床,但高性能模型可能没有以前想象的那么多多跳的推理。
Learning multi-hop reasoning has been a key challenge for reading comprehension models, leading to the design of datasets that explicitly focus on it. Ideally, a model should not be able to perform well on a multi-hop question answering task without doing multi-hop reasoning. In this paper, we investigate two recently proposed datasets, WikiHop and HotpotQA. First, we explore sentence-factored models for these tasks; by design, these models cannot do multi-hop reasoning, but they are still able to solve a large number of examples in both datasets. Furthermore, we find spurious correlations in the unmasked version of WikiHop, which make it easy to achieve high performance considering only the questions and answers. Finally, we investigate one key difference between these datasets, namely span-based vs. multiple-choice formulations of the QA task. Multiple-choice versions of both datasets can be easily gamed, and two models we examine only marginally exceed a baseline in this setting. Overall, while these datasets are useful testbeds, high-performing models may not be learning as much multi-hop reasoning as previously thought.