A Probabilistic Annotation Model for Crowdsourcing Coreference

A Probabilistic Annotation Model for Crowdsourcing Coreference
复制标题

众包共指的概率注释模型

DOI:
10.18653/v1/d18-1218
复制
发表时间:
2018
期刊:
ACM Trans. Interact. Intell. Syst.
影响因子:
--
通讯作者:
Massimo Poesio
Massimo Poesio
中科院分区:
--
文献类型:
--
作者:
Silviu Paun;Jon Chamberlain;Udo Kruschwitz;Juntao Yu;Massimo Poesio

文献摘要

被引文献

相似文献

大规模注释语料库的可供共指对该领域的发展至关重要。然而,通过专家注释以所需规模创建资源将过于昂贵。众包已被提议作为一种替代办法,但这种办法尚未被广泛用于共同参考。本文解决了一个关键的障碍的方式,使这成为可能,通过引入一个新的模型的注释聚合众包照应注释。该模型沿沿着三个维度进行评估:推断的提及对的准确性,事后构建的银链的质量,以及使用银链作为训练最先进的共指系统中专家注释链的替代品的可行性。结果表明,我们的模型可以从众包注释中提取质量与专家注释相当的共指链。
The availability of large scale annotated corpora for coreference is essential to the development of the field. However, creating resources at the required scale via expert annotation would be too expensive. Crowdsourcing has been proposed as an alternative; but this approach has not been widely used for coreference. This paper addresses one crucial hurdle on the way to make this possible, by introducing a new model of annotation for aggregating crowdsourced anaphoric annotations. The model is evaluated along three dimensions: the accuracy of the inferred mention pairs, the quality of the post-hoc constructed silver chains, and the viability of using the silver chains as an alternative to the expert-annotated chains in training a state of the art coreference system. The results suggest that our model can extract from crowdsourced annotations coreference chains of comparable quality to those obtained with expert annotation.