Building Large-Scale Twitter-Specific Sentiment Lexicon : A Representation Learning Approach

Building Large-Scale Twitter-Specific Sentiment Lexicon : A Representation Learning Approach
复制标题

DOI:
--
复制
发表时间:
2014-08
期刊:
--
影响因子:
--
通讯作者:
Duyu Tang;Furu Wei;Bing Qin;M. Zhou;Ting Liu
Duyu Tang;Furu Wei;Bing Qin;M. Zhou;Ting Liu
中科院分区:
其他
文献类型:
--
作者:
Duyu Tang;Furu Wei;Bing Qin;M. Zhou;Ting Liu

文献摘要

被引文献

相似文献

在本文中,我们建议使用表示学习方法从Twitter中构建大规模情感词典。我们将情感词典学习作为短语级情感分类任务。面临的挑战是开发有效的短语特征表示,并获得训练数据与轻微的手动注释,以建立情感分类器。具体来说,我们开发了一个专用的神经架构,并将文本(例如句子或推文)的情感信息集成到其混合损失函数中,用于学习特定于情感的短语嵌入(SSPE)。神经网络是从大量的tweet中训练出来的,这些tweet中包含了积极和消极的表情符号,没有任何手动注释。此外,我们引入城市词典来扩展少量的情感种子,以获得更多的训练数据来构建短语级情感分类器。我们评估我们的情感词典(TS-Lex)通过将其应用于Twitter情感分类的监督学习框架。在SemEval 2013的基准数据集上的实验结果表明,TS-Lex比以前引入的情感词典具有更好的性能。
In this paper, we propose to build large-scale sentiment lexicon from Twitter with a representation learning approach. We cast sentiment lexicon learning as a phrase-level sentiment classification task. The challenges are developing effective feature representation of phrases and obtaining training data with minor manual annotations for building the sentiment classifier. Specifically, we develop a dedicated neural architecture and integrate the sentiment information of text (e.g. sentences or tweets) into its hybrid loss function for learning sentiment-specific phrase embedding (SSPE). The neural network is trained from massive tweets collected with positive and negative emoticons, without any manual annotation. Furthermore, we introduce the Urban Dictionary to expand a small number of sentiment seeds to obtain more training data for building the phrase-level sentiment classifier. We evaluate our sentiment lexicon (TS-Lex) by applying it in a supervised learning framework for Twitter sentiment classification. Experiment results on the benchmark dataset of SemEval 2013 show that, TS-Lex yields better performance than previously introduced sentiment lexicons.