Constructing Corpora for the Development and Evaluation of Paraphrase Systems

Constructing Corpora for the Development and Evaluation of Paraphrase Systems
复制标题

DOI:
10.1162/coli.08-003-r1-07-044
复制
发表时间:
2008-12-01
影响因子:
9.3
通讯作者:
Lapata, Mirella
Lapata, Mirella
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cohn, Trevor;Callison-Burch, Chris;Lapata, Mirella

文献摘要

被引文献

相似文献

自动释义是许多自然语言处理任务中的重要组成部分。在本文中,我们介绍了带有释义注释的新的平行语料库。我们基于单词对齐方式采用了释义的定义,并表明它产生了高通道的一致性。由于Kappa适合名义数据,因此我们采用了一个替代协议统计量,适用于结构化的一致性任务。我们讨论如何自动评估术语系统(例如,通过测量精度,召回和F1)以及基于句法结构开发语言富含术语的释义模型。
Automatic paraphrasing is an important component in many natural language processing tasks. In this article we present a new parallel corpus with paraphrase annotations. We adopt a definition of paraphrase based on word alignments and show that it yields high inter-annotator agreement. As Kappa is suited to nominal data, we employ an alternative agreement statistic which is appropriate for structured alignment tasks. We discuss how the corpus can be usefully employed in evaluating paraphrase systems automatically ( e. g., by measuring precision, recall, and F1) and also in developing linguistically rich paraphrase models based on syntactic structure.