ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic Relations

ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic Relations
复制标题

DOI:
10.18653/v1/2021.emnlp-main.597
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Rujun Han;I-Hung Hsu;Jiao Sun;J. Baylón;Qiang Ning;D. Roth;Nanyun Peng
Rujun Han;I-Hung Hsu;Jiao Sun;J. Baylón;Qiang Ning;D. Roth;Nanyun Peng
中科院分区:
其他
文献类型:
--
作者:
Rujun Han;I-Hung Hsu;Jiao Sun;J. Baylón;Qiang Ning;D. Roth;Nanyun Peng

文献摘要

被引文献

相似文献

理解事件之间的语义关系是阅读理解的本质。最近以事件为中心的阅读理解数据集主要关注事件参数或时间关系。虽然这些任务部分评估了机器的叙事理解能力,但类似人类的阅读理解需要机器处理基于事件的信息的能力,而不是争论和时间推理。例如,为了理解事件之间的因果关系,我们需要推断动机或目的;要建立事件层次结构,我们需要了解事件的组成。为了方便这些任务,我们引入了**ESTER**,一个用于事件语义关系推理的综合机器阅读理解(MRC)数据集。该数据集利用自然语言查询来推断五种最常见的事件语义关系,提供了6K多个问题,并捕获了10.1K个事件关系对。实验结果表明,当前SOTA系统在基于令牌的精确匹配(**EM**)、**F1**和基于事件的**HIT@1**得分分别达到22.1%、63.3%和83.5%,均显著低于人类的表现(分别为36.0%、79.6%和100%),突出了我们的数据集是一个具有挑战性的基准。
Understanding how events are semantically related to each other is the essence of reading comprehension. Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations. While these tasks partially evaluate machines’ ability of narrative understanding, human-like reading comprehension requires the capability to process event-based information beyond arguments and temporal reasoning. For example, to understand causality between events, we need to infer motivation or purpose; to establish event hierarchy, we need to understand the composition of events. To facilitate these tasks, we introduce **ESTER**, a comprehensive machine reading comprehension (MRC) dataset for Event Semantic Relation Reasoning. The dataset leverages natural language queries to reason about the five most common event semantic relations, provides more than 6K questions, and captures 10.1K event relation pairs. Experimental results show that the current SOTA systems achieve 22.1%, 63.3% and 83.5% for token-based exact-match (**EM**), **F1** and event-based **HIT@1** scores, which are all significantly below human performances (36.0%, 79.6%, 100% respectively), highlighting our dataset as a challenging benchmark.