Removing the Training Wheels: A Coreference Dataset that Entertains Humans and Challenges Computers

Removing the Training Wheels: A Coreference Dataset that Entertains Humans and Challenges Computers
复制标题

去掉辅助轮:一个既娱乐人类又挑战计算机的共指数据集

DOI:
--
复制
发表时间:
2015
期刊:
North American Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Jordan L. Boyd
Jordan L. Boyd
中科院分区:
--
文献类型:
--
作者:
Anupam Guha;Mohit Iyyer;D. Bouman;Jordan L. Boyd

文献摘要

被引文献

相似文献

共指是NLP的核心问题。然而,作为现有共引数据的主要来源,新闻通讯社的数据缺乏真正解决共引所需的丰富性。我们提出了一个具有更密集引用的新领域-测验碗问题-对人类来说既具有挑战性又很有趣,我们使用测验碗社区开发了一个新的共引用数据集,以及一个可以用共引用和命名实体标记任何文本数据的注释框架。我们还成功地将主动学习集成到这个注释管道中,以收集对相互引用模型最有用的文档。最先进的共参考文献系统在我们的新数据集上的表现逊于简单的分类器,激励了非新闻网站的数据用于未来的共参考文献研究。
Coreference is a core nlp problem. However, newswire data, the primary source of existing coreference data, lack the richness necessary to truly solve coreference. We present a new domain with denser references—quiz bowl questions—that is challenging and enjoyable to humans, and we use the quiz bowl community to develop a new coreference dataset, together with an annotation framework that can tag any text data with coreferences and named entities. We also successfully integrate active learning into this annotation pipeline to collect documents maximally useful to coreference models. State-of-the-art coreference systems underperform a simple classifier on our new dataset, motivating non-newswire data for future coreference research.