Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora.
Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora.
复制标题
DOI:
10.18653/v1/d16-1057
复制
发表时间:
2016-11
期刊:
影响因子:
--
通讯作者:
Jurafsky D
中科院分区:
文献类型:
--
作者:
Hamilton WL;Clark K;Leskovec J;Jurafsky D
A word’s sentiment depends on the domain in which it is used. Computational social science research thus requires sentiment lexicons that are specific to the domains being studied. We combine domain-specific word embeddings with a label propagation framework to induce accurate domain-specific sentiment lexicons using small sets of seed words. We show that our approach achieves state-of-the-art performance on inducing sentiment lexicons from domain-specific corpora and that our purely corpus-based approach outperforms methods that rely on hand-curated resources (e.g., WordNet). Using our framework, we induce and release historical sentiment lexicons for 150 years of English and community-specific sentiment lexicons for 250 online communities from the social media forum Reddit. The historical lexicons we induce show that more than 5% of sentiment-bearing (non-neutral) English words completely switched polarity during the last 150 years, and the community-specific lexicons highlight how sentiment varies drastically between different communities.