Analysis of named entity recognition and linking for tweets

Analysis of named entity recognition and linking for tweets
复制标题

DOI:
10.1016/j.ipm.2014.10.006
复制
发表时间:
2015-03-01
影响因子:
8.6
通讯作者:
Bontcheva, Kalina
Bontcheva, Kalina
中科院分区:
计算机科学1区
文献类型:
--
作者:
Derczynski, Leon;Maynard, Diana;Bontcheva, Kalina

文献摘要

被引文献

相似文献

将自然语言处理应用于挖掘推文和智能信息访问(微博的一种形式)是一个具有挑战性的新兴研究领域。与精心撰写的新闻文本和其他较长的内容不同,推文因其简短、嘈杂、上下文相关和动态的性质而提出了许多新的挑战。从推文中提取信息通常是在流水线中执行的,包括语言识别、标记化、词性标记、命名实体识别和实体歧义消除(例如关于DBpedia)的连续阶段。在这项工作中,我们描述了一个新的Twitter实体消歧数据集,并对命名实体识别和消歧进行了实证分析,调查了一些最先进的系统在这样的噪声文本上的健壮性,主要的错误来源是什么,以及哪些问题需要进一步研究以提高技术水平。(C)爱思唯尔有限公司出版的2015年。
Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number of new challenges, due to their short, noisy, context-dependent, and dynamic nature. Information extraction from tweets is typically performed in a pipeline, comprising consecutive stages of language identification, tokenisation, part-of-speech tagging, named entity recognition and entity disambiguation (e.g. with respect to DBpedia). In this work, we describe a new Twitter entity disambiguation dataset, and conduct an empirical analysis of named entity recognition and disambiguation, investigating how robust a number of state-of-the-art systems are on such noisy texts, what the main sources of error are, and which problems should be further investigated to improve the state of the art. (C) 2015 Published by Elsevier Ltd.