Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora.

Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora.
复制标题

DOI:
10.18653/v1/d16-1057
复制
发表时间:
2016-11
期刊:
Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Jurafsky D
Jurafsky D
中科院分区:
其他
文献类型:
--
作者:
Hamilton WL;Clark K;Leskovec J;Jurafsky D

文献摘要

被引文献

相似文献

一个词的情感取决于它被使用的领域。因此,计算社会科学研究需要特定于所研究领域的情感词典。我们结合联合收割机特定领域的词嵌入与标签传播框架,以诱导准确的特定领域的情感词典使用小的种子词集。我们表明,我们的方法在从特定领域的语料库中诱导情感词典方面达到了最先进的性能,并且我们纯粹基于语料库的方法优于依赖于手工策划资源的方法(例如,WordNet)。使用我们的框架,我们诱导和发布历史情感词汇150年的英语和社区特定的情感词汇250个在线社区从社交媒体论坛Reddit。我们归纳的历史词汇表明,在过去的150年中,超过5%的情感承载(非中性)英语单词完全转换了极性,而社区特定词汇则突出了不同社区之间情感的巨大差异。
A word’s sentiment depends on the domain in which it is used. Computational social science research thus requires sentiment lexicons that are specific to the domains being studied. We combine domain-specific word embeddings with a label propagation framework to induce accurate domain-specific sentiment lexicons using small sets of seed words. We show that our approach achieves state-of-the-art performance on inducing sentiment lexicons from domain-specific corpora and that our purely corpus-based approach outperforms methods that rely on hand-curated resources (e.g., WordNet). Using our framework, we induce and release historical sentiment lexicons for 150 years of English and community-specific sentiment lexicons for 250 online communities from the social media forum Reddit. The historical lexicons we induce show that more than 5% of sentiment-bearing (non-neutral) English words completely switched polarity during the last 150 years, and the community-specific lexicons highlight how sentiment varies drastically between different communities.