TwitIE: An Open-Source Information Extraction Pipeline for Microblog Text

TwitIE: An Open-Source Information Extraction Pipeline for Microblog Text
复制标题

DOI:
10.6084/m9.figshare.1003767.v2
复制
发表时间:
2013-09
期刊:
--
影响因子:
--
通讯作者:
Kalina Bontcheva;Leon Derczynski;Adam Funk;M. Greenwood;D. Maynard;N. Aswani
Kalina Bontcheva;Leon Derczynski;Adam Funk;M. Greenwood;D. Maynard;N. Aswani
中科院分区:
其他
文献类型:
--
作者:
Kalina Bontcheva;Leon Derczynski;Adam Funk;M. Greenwood;D. Maynard;N. Aswani

文献摘要

相似文献

Twitter是微博文本的最大来源,每天负责千兆字节的人类话语。处理微博文本比较困难:体裁嘈杂,文档缺乏上下文,话语非常短。因此,传统的NLP工具在面对tweet和其他微博文本时失败了。我们提出了twitter,一个开源的NLP管道,在每个阶段定制微博文本。此外,它还包括特定于twitter的数据导入和元数据处理。本文介绍了tweetie管道的各个阶段,tweetie管道是对GATE ANNIE开源新闻文本管道的改进。并对一些最先进的系统进行了评价。
Twitter is the largest source of microblog text, responsible for gigabytes of human discourse every day. Processing microblog text is difficult: the genre is noisy, documents have little context, and utterances are very short. As such, conventional NLP tools fail when faced with tweets and other microblog text. We present TwitIE, an open-source NLP pipeline customised to microblog text at every stage. Additionally, it includes Twitter-specific data import and metadata handling. This paper introduces each stage of the TwitIE pipeline, which is a modification of the GATE ANNIE open-source pipeline for news text. An evaluation against some state-of-the-art systems is also presented.