Removing the Training Wheels: A Coreference Dataset that Entertains Humans and Challenges Computers
Removing the Training Wheels: A Coreference Dataset that Entertains Humans and Challenges Computers
复制标题
去掉辅助轮:一个既娱乐人类又挑战计算机的共指数据集
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Jordan L. Boyd
中科院分区:
文献类型:
--
作者:
Anupam Guha;Mohit Iyyer;D. Bouman;Jordan L. Boyd
Coreference is a core nlp problem. However, newswire data, the primary source of existing coreference data, lack the richness necessary to truly solve coreference. We present a new domain with denser references—quiz bowl questions—that is challenging and enjoyable to humans, and we use the quiz bowl community to develop a new coreference dataset, together with an annotation framework that can tag any text data with coreferences and named entities. We also successfully integrate active learning into this annotation pipeline to collect documents maximally useful to coreference models. State-of-the-art coreference systems underperform a simple classifier on our new dataset, motivating non-newswire data for future coreference research.