The Unreasonable Effectiveness of Word Representations for Twitter Named Entity Recognition

The Unreasonable Effectiveness of Word Representations for Twitter Named Entity Recognition
复制标题

DOI:
10.3115/v1/n15-1075
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
Colin Cherry;Hongyu Guo
Colin Cherry;Hongyu Guo
中科院分区:
其他
文献类型:
--
作者:
Colin Cherry;Hongyu Guo

文献摘要

被引文献

相似文献

在Twitter上测试时,在新闻专线上训练的命名实体识别(NER)系统的表现非常糟糕。在编辑文本中可靠的信号几乎完全消失在Twitter的非正式聊天中,需要构建专门的模型。使用众所周知的技术,我们开始提高Twitter NER性能时,给出了一个小的一组注释的训练推文。为了利用未标记的推文,我们构建了布朗聚类和词向量,从而实现了分布相似词的泛化。为了利用带注释的新闻专线数据,我们采用了重要性加权方案。综上所述,我们建立了一个新的国家的最先进的两个共同的测试集。虽然众所周知,词表征是有用的NER,支持实验迄今为止集中在新闻专线数据。我们强调Twitter NER上的表示的有效性,并证明他们的列入可以提高性能高达20 F1。
Named entity recognition (NER) systems trained on newswire perform very badly when tested on Twitter. Signals that were reliable in copy-edited text disappear almost entirely in Twitter’s informal chatter, requiring the construction of specialized models. Using wellunderstood techniques, we set out to improve Twitter NER performance when given a small set of annotated training tweets. To leverage unlabeled tweets, we build Brown clusters and word vectors, enabling generalizations across distributionally similar words. To leverage annotated newswire data, we employ an importance weighting scheme. Taken all together, we establish a new state-of-the-art on two common test sets. Though it is wellknown that word representations are useful for NER, supporting experiments have thus far focused on newswire data. We emphasize the effectiveness of representations on Twitter NER, and demonstrate that their inclusion can improve performance by up to 20 F1.