TwitIE: An Open-Source Information Extraction Pipeline for Microblog Text
TwitIE: An Open-Source Information Extraction Pipeline for Microblog Text
复制标题
DOI:
10.6084/m9.figshare.1003767.v2
复制
发表时间:
2013-09
期刊:
影响因子:
--
通讯作者:
Kalina Bontcheva;Leon Derczynski;Adam Funk;M. Greenwood;D. Maynard;N. Aswani
中科院分区:
文献类型:
--
作者:
Kalina Bontcheva;Leon Derczynski;Adam Funk;M. Greenwood;D. Maynard;N. Aswani
Twitter is the largest source of microblog text, responsible for gigabytes of human discourse every day. Processing microblog text is difficult: the genre is noisy, documents have little context, and utterances are very short. As such, conventional NLP tools fail when faced with tweets and other microblog text. We present TwitIE, an open-source NLP pipeline customised to microblog text at every stage. Additionally, it includes Twitter-specific data import and metadata handling. This paper introduces each stage of the TwitIE pipeline, which is a modification of the GATE ANNIE open-source pipeline for news text. An evaluation against some state-of-the-art systems is also presented.