Behind the Scenes of an Evolving Event Cloze Test

Behind the Scenes of an Evolving Event Cloze Test
复制标题

演变事件完形填空测试的幕后故事

DOI:
--
复制
发表时间:
2017
期刊:
LSDSem@EACL
影响因子:
--
通讯作者:
Nathanael Chambers
Nathanael Chambers
中科院分区:
--
文献类型:
--
作者:
Nathanael Chambers

文献摘要

被引文献

相似文献

本文分析了叙事事件完形填空测试及其近年来的演变。测试从文档的事件链中删除一个事件,然后系统预测缺失的事件。最初提议评估事件场景(例如,脚本和框架)的学习知识,最近的工作现在构建了类语法的语言模型(LM)来通过测试。本文认为,测试已经慢慢地/不知不觉地被改变以适应LMs.5最值得注意的是,测试是自动生成的,而不是手工生成的,并且不需要花费精力来包含核心脚本事件。最近的工作对评价目标不明确,结果相互矛盾。我们实现了几个模型,并表明测试对高频事件的偏差解释了不一致性。最后,我们就如何回归测试的初衷提出了建议,并就前进的道路提出了简短的建议。
This paper analyzes the narrative event cloze test and its recent evolution. The test removes one event from a document’s chain of events, and systems predict the missing event. Originally proposed to evaluate learned knowledge of event scenarios (e.g., scripts and frames), most recent work now builds ngram-like language models (LM) to beat the test. This paper argues that the test has slowly/unknowingly been altered to accommodate LMs.5 Most notably, tests are auto-generated rather than by hand, and no effort is taken to include core script events. Recent work is not clear on evaluation goals and contains contradictory results. We implement several models, and show that the test’s bias to high-frequency events explains the inconsistencies. We conclude with recommendations on how to return to the test’s original intent, and offer brief suggestions on a path forward.