A Dataset for Tracking Entities in Open Domain Procedural Text

A Dataset for Tracking Entities in Open Domain Procedural Text
复制标题

用于跟踪开放域程序文本中的实体的数据集

DOI:
--
复制
发表时间:
2020
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
E. Hovy
E. Hovy
中科院分区:
--
文献类型:
--
作者:
Niket Tandon;Keisuke Sakaguchi;Bhavana Dalvi;Dheeraj Rajagopal;Peter Clark;Michal Guerquin;Kyle Richardson;E. Hovy

文献摘要

被引文献

相似文献

我们提出了第一个数据集,用于通过使用不受限制(开放)的词汇来跟踪来自任意域的程序文本的状态变化。例如,在描述使用土豆除雾的文本中,车窗可能在有雾、粘性、不透明和透明之间转变。此任务的先前表述提供了所涉及的文本和实体,并询问这些实体如何针对一小部分预定义的属性(例如位置)进行更改,从而限制了它们的保真度。我们的解决方案是一个新的任务公式,其中仅给定程序文本作为输入,任务是为每个步骤生成一组状态更改元组(实体、属性、前状态、后状态),其中实体、属性和状态值必须从开放词汇表中预测。通过众包,我们创建了 OPENPI,这是一个高质量(由人类判断并经过全面审查的覆盖率为 91.5%)的大型数据集,包含来自 WikiHow.com 的 810 个程序现实世界段落中的 4,050 个句子中的 29,928 个状态变化。基于 BLEU 指标,当前最先进的生成模型在该任务上实现了 16.1% F1,为新颖的模型架构留下了足够的空间。
We present the first dataset for tracking state changes in procedural text from arbitrary domains by using an unrestricted (open) vocabulary. For example, in a text describing fog removal using potatoes, a car window may transition between being foggy, sticky, opaque, and clear. Previous formulations of this task provide the text and entities involved, and ask how those entities change for just a small, pre-defined set of attributes (e.g., location), limiting their fidelity. Our solution is a new task formulation where given just a procedural text as input, the task is to generate a set of state change tuples (entity, attribute, before-state, after-state) for each step, where the entity, attribute, and state values must be predicted from an open vocabulary. Using crowdsourcing, we create OPENPI, a high-quality (91.5% coverage as judged by humans and completely vetted), and large-scale dataset comprising 29,928 state changes over 4,050 sentences from 810 procedural real-world paragraphs from WikiHow.com. A current state-of-the-art generation model on this task achieves 16.1% F1 based on BLEU metric, leaving enough room for novel model architectures.