Generation of repeated references to discourse entities

Generation of repeated references to discourse entities
复制标题

生成对话语实体的重复引用

DOI:
10.3115/1610163.1610167
复制
发表时间:
2007
影响因子:
1.3
通讯作者:
S. Varges
S. Varges
中科院分区:
医学4区
文献类型:
--
作者:
A. Belz;S. Varges

文献摘要

被引文献

相似文献

引用表达式的生成是自然语言生成的一个蓬勃发展的子领域,传统上它专注于选择一组明确标识给定引用对象的属性的任务。在本文中,我们解决了生成重复的、可能不同的指称表达的互补问题,这些表达在一段比句子长的话语中指代同一实体。我们描述了我们编译和注释的简短百科全书文本语料库,以参考文本的主要主题,并报告我们的实验结果,其中我们将人类受试者和自动方法设置为从全文上下文中的广泛选择中选择参考表达的任务。我们发现,我们的人类受试者在表达选择上有相当大的程度的一致,在 50% 的情况下选择了三个相同的表达。我们测试了基于最常见选择启发式的自动选择策略,涉及有关句法 MSR 类型和域类型的信息的不同组合。我们发现,更多的信息通常会产生更好的结果,当句法 MSR 类型和域类型都已知时,总体测试集准确率达到 53.9%。
Generation of Referring Expressions is a thriving subfield of Natural Language Generation which has traditionally focused on the task of selecting a set of attributes that unambiguously identify a given referent. In this paper, we address the complementary problem of generating repeated, potentially different referential expressions that refer to the same entity in the context of a piece of discourse longer than a sentence. We describe a corpus of short encyclopaedic texts we have compiled and annotated for reference to the main subject of the text, and report results for our experiments in which we set human subjects and automatic methods the task of selecting a referential expression from a wide range of choices in a full-text context. We find that our human subjects agree on choice of expression to a considerable degree, with three identical expressions selected in 50% of cases. We tested automatic selection strategies based on most frequent choice heuristics, involving different combinations of information about syntactic MSR type and domain type. We find that more information generally produces better results, achieving a best overall test set accuracy of 53.9% when both syntactic MSR type and domain type are known.