Computing Reliability for Coreference Annotation

Computing Reliability for Coreference Annotation
复制标题

共指注释的可靠性计算

DOI:
10.7916/d8w09fct
复制
发表时间:
2004
期刊:
Knowl. Based Syst.
影响因子:
--
通讯作者:
R. Passonneau
R. Passonneau
中科院分区:
--
文献类型:
--
作者:
R. Passonneau

文献摘要

被引文献

相似文献

共指标注是对语言语料库进行标注,指出哪些表达方式被用来共指同一语篇实体。当从两个或多个编码器收集相同数据的注释时,可能需要对数据的可靠性进行量化。在应用可靠性度量的过程中存在两个障碍:跨注释的不相称单位,以及缺乏编码值的方便表示。给定N个编码器和M个编码单元,可靠性根据N x M矩阵计算,该矩阵记录编码器N k分配给单元M j的值。我提出的解决方案为注释器提供了广泛的编码选择,同时在编码中保持相同的单位。因此,它允许一个简单的应用程序的可靠性测量。此外,在共指标注中,分歧可以是完全的或部分的,所以我引入了一个距离度量来衡量分歧。这种方法也被应用到一个相当独特的编码任务,即语义注释的摘要。
Coreference annotation is annotation of language corpora to indicate which expressions have been used to co-specify the same discourse entity. When annotations of the same data are collected from two or more coders, the reliability of the data may need to be quanti(cid:2)ed. Two obstacles have stood in the way of applying reliability metrics: incommensurate units across annotations, and lack of a convenient representation of the coding values. Given N coders and M coding units, reliability is computed from an N-by-M matrix that records the value assigned to unit M j by coder N k . The solution I present accommodates a wide range of coding choices for the annotator, while preserving the same units across codings. As a consequence, it permits a straightforward application of reliability measurement. In addition, in coreference annotation, disagreements can be complete or partial so I incorporate a distance metric to scale disagreements. This method has also been applied to a quite distinct coding task, namely semantic annotation of summaries.