Analysis : It ’ s Complicated !

Analysis : It ’ s Complicated !
复制标题

分析:很复杂!

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
D. Ruths
D. Ruths
中科院分区:
--
文献类型:
--
作者:
Kian Kenyon;Eisha Ahmed;Scott Fujimoto;Jeremy Georges;Christopher Glasz;Barleen Kaur;Auguste Lalande;Shruti Bhanderi;Robert Belfer;N. Kanagasabai;Roman Sarrazingendron;Rohit Verma;D. Ruths

文献摘要

被引文献

相似文献

情感分析用作衡量人类情感的代理,其目标是根据某种预定义的情感概念对文本进行分类。情感分析数据集通常使用黄金标准情感标签构建,并根据手动注释的结果进行分配。在使用此类注释时,数据集构造者通常会丢弃“嘈杂”或“有争议”的数据,因为这些数据在正确的标签上存在重大分歧。在为 Twitter 情感分析 (TSA) 构建的数据集中,这些有争议的示例可能占原始注释数据的 30% 以上。我们认为,删除此类数据是一个有问题的趋势,因为在对短文本进行实时情感分类时,自动化系统无法先验地知道哪些样本会属于此类有争议的情感。因此,我们提出“复杂”情感类的概念来对此类文本进行分类,并认为将其包含在短文本情感分析框架中将提高自动情感分析系统在现实世界环境中实施时的质量。我们通过构建和分析一个新的公开可用的 TSA 数据集(名为 MTSA)来激发这一论点,该数据集包含 7,000 多条推文,并标注了 5 倍覆盖率。我们对数据集上的分类器性能的分析提供了对情感分析数据集和模型设计、当前技术在现实世界中如何执行以及研究人员应如何处理困难数据的见解。
Sentiment analysis is used as a proxy to measure human emotion, where the objective is to categorize text according to some predefined notion of sentiment. Sentiment analysis datasets are typically constructed with gold-standard sentiment labels, assigned based on the results of manual annotations. When working with such annotations, it is common for dataset constructors to discard “noisy” or “controversial” data where there is significant disagreement on the proper label. In datasets constructed for the purpose of Twitter sentiment analysis (TSA), these controversial examples can compose over 30% of the originally annotated data. We argue that the removal of such data is a problematic trend because, when performing real-time sentiment classification of short-text, an automated system cannot know a priori which samples would fall into this category of disputed sentiment. We therefore propose the notion of a “complicated” class of sentiment to categorize such text, and argue that its inclusion in the short-text sentiment analysis framework will improve the quality of automated sentiment analysis systems as they are implemented in real-world settings. We motivate this argument by building and analyzing a new publicly available TSA dataset of over 7,000 tweets annotated with 5x coverage, named MTSA. Our analysis of classifier performance over our dataset offers insights into sentiment analysis dataset and model design, how current techniques would perform in the real world, and how researchers should handle difficult data.