Refining Word Embeddings Using Intensity Scores for Sentiment Analysis

Refining Word Embeddings Using Intensity Scores for Sentiment Analysis
复制标题

使用强度分数细化词嵌入进行情感分析

DOI:
10.1109/taslp.2017.2788182
复制
发表时间:
2018-03-01
影响因子:
5.4
通讯作者:
Zhang, Xuejie
Zhang, Xuejie
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yu, Liang-Chih;Wang, Jin;Zhang, Xuejie

文献摘要

被引文献

相似文献

提供词的连续低维向量表示的词嵌入已被广泛用于各种自然语言处理任务。然而,诸如Word2vec和GloVe之类的现有的基于上下文的词嵌入通常无法捕获足够的情感信息,这可能导致具有类似向量表示的词具有相反的情感极性(例如,好的和坏的),从而降低情感分析性能。为了解决这个问题,最近的研究建议学习情感嵌入,以纳入来自标记语料库的情感极性(积极和消极)信息。本研究采用另一种策略来学习情感嵌入。我们提出了一个词向量细化模型,使用情感词典提供的实值情感强度分数来细化现有的预训练词向量,而不是从标记的语料库中创建一个新的词嵌入。细化模型的思想是改进每个词向量,使得它在词典中可以更接近语义和情感上相似的词(即,具有相似强度分数的那些)并且远离情感上不同的词(即,具有不同强度分数的那些)。该方法的一个明显优点是它可以应用于任何预训练的词嵌入。此外,强度分数可以提供比二进制极性标签更细粒度(实值)的情感信息,以指导细化过程。在SemEval和斯坦福大学情感树库数据集上进行的实验结果表明,该改进模型可以改进传统的词嵌入和先前提出的二进制、三进制和细粒度情感分类的情感嵌入。
Word embeddings that provide continuous lowdimensional vector representations of words have been extensively used for various natural language processing tasks. However, existing context-based word embeddings such as Word2vec and GloVe typically fail to capture sufficient sentiment information, which may result in words with similar vector representations having an opposite sentiment polarity (e.g., good and bad), thus degrading sentiment analysis performance. To tackle this problem, recent studies have suggested learning sentiment embeddings to incorporate the sentiment polarity (positive and negative) information from labeled corpora. This study adopts another strategy to learn sentiment embeddings. Instead of creating a new word embedding from labeled corpora, we propose a word vector refinement model to refine existing pretrained word vectors using real-valued sentiment intensity scores provided by sentiment lexicons. The idea of the refinement model is to improve each word vector such that it can be closer in the lexicon to both semantically and sentimentally similar words (i.e., those with similar intensity scores) and further away from sentimentally dissimilar words (i.e., those with dissimilar intensity scores). An obvious advantage of the proposed method is that it can be applied to any pretrained word embeddings. In addition, the intensity scores can provide more finegrained (real-valued) sentiment information than binary polarity labels to guide the refinement process. Experimental results show that the proposed refinement model can improve both conventional word embeddings and previously proposed sentiment embeddings for binary, ternary, and fine- grained sentiment classification on the SemEval and Stanford Sentiment Treebank datasets.