Learning Sentence Representation for Emotion Classification on Microblogs

Learning Sentence Representation for Emotion Classification on Microblogs
复制标题

DOI:
10.1007/978-3-642-41644-6_20
复制
发表时间:
2013-11
期刊:
--
影响因子:
--
通讯作者:
Duyu Tang;Bing Qin;Ting Liu;Zhenghua Li
Duyu Tang;Bing Qin;Ting Liu;Zhenghua Li
中科院分区:
其他
文献类型:
--
作者:
Duyu Tang;Bing Qin;Ting Liu;Zhenghua Li

文献摘要

被引文献

相似文献

本文研究了微博上的情感分类任务。给定一条消息,我们将其情绪分为快乐,悲伤,愤怒或惊讶。现有的方法大多使用词袋表示或手动设计的特征来训练监督或远程监督模型。然而,制造特征引擎是耗时的,不足以捕捉微博上复杂的语言现象。在本研究中,为了克服上述问题,我们利用在Twitter情感分析中广泛用于远程监督学习和训练语言模型的伪标记数据,通过深度信念网络算法学习句子表示。在监督学习框架中的实验结果表明,使用伪标记数据,通过深度信念网络学习的表示优于基于主成分分析和基于潜在狄利克雷分配的表示。通过将基于深度信念网络的表示纳入基本特征,性能得到进一步提高。
This paper studies the emotion classification task on microblogs. Given a message, we classify its emotion as happy, sad, angry or surprise. Existing methods mostly use the bag-of-word representation or manually designed features to train supervised or distant supervision models. However, manufacturing feature engines is time-consuming and not enough to capture the complex linguistic phenomena on microblogs. In this study, to overcome the above problems, we utilize pseudo-labeled data, which is extensively explored for distant supervision learning and training language model in Twitter sentiment analysis, to learn the sentence representation through Deep Belief Network algorithm. Experimental results in the supervised learning framework show that using the pseudo-labeled data, the representation learned by Deep Belief Network outperforms the Principal Components Analysis based and Latent Dirichlet Allocation based representations. By incorporating the Deep Belief Network based representation into basic features, the performance is further improved.