Sentiment analysis on big sparse data streams with limited labels

Sentiment analysis on big sparse data streams with limited labels
复制标题

DOI:
10.1007/s10115-019-01392-9
复制
发表时间:
2020-04-01
影响因子:
2.7
通讯作者:
Ntoutsi, Eirini
Ntoutsi, Eirini
中科院分区:
计算机科学4区
文献类型:
--
作者:
Iosifidis, Vasileios;Ntoutsi, Eirini

文献摘要

被引文献

相似文献

情绪分析是一项重要的任务,以便深入了解Twitter等社交媒体每天产生的大量固执己见的文本。尽管数量巨大,但标准的监督学习方法不会对这类数据起作用,因为缺乏标签,而且在这种规模下(人类)标记是不切实际的。在这项工作中,我们利用远程监督和半监督学习来注释2015年的大量推文,其中包括2.28亿条没有转发的推文(以及2.75亿条有转发的推文)。我们提出了注释过程中关于不同半监督学习方法(即自学习、联合训练和期望最大化)效果的见解。此外,我们提出了两种注释模式,批处理模式,其中所有标记和未标记的数据从一开始就可用于算法和轻量级流模式,该模式根据数据在流中的到达时间分批处理数据。我们的实验表明,流处理与三个月的滑动窗口实现了可比的结果,批处理,同时更有效。最后,为了解决类不平衡问题,由于我们的数据集对积极情绪类不平衡,以及半监督学习方法的加剧,我们在半监督学习过程中采用数据增强,以均衡类分布。我们的研究结果表明,半监督学习与数据增强相结合,显着优于默认的半监督注释过程。我们向社区提供所谓的TSentiment15情感注释数据集,用于评估目的和开发新方法。
Sentiment analysis is an important task in order to gain insights over the huge amounts of opinionated texts generated on a daily basis in social media like Twitter. Despite its huge amount, standard supervised learning methods won't work upon such sort of data due to lack of labels and the impracticality of (human) labeling at this scale. In this work, we leverage distant supervision and semi-supervised learning to annotate a big stream of tweets from 2015 which consists of 228 million tweets without retweets (and 275 million with retweets). We present the insights from our annotation process regarding the effect of different semi-supervised learning approaches, namely Self-Learning, Co-Training and Expectation-Maximization. Moreover, we propose two annotation modes, the batch mode where all labeled and unlabeled data are available to the algorithms from the beginning and a lightweight streaming mode that processes the data in batches based on their arrival time in the stream. Our experiments show that stream processing with a sliding window of three months achieves comparable results to batch processing while being more efficient. Finally, to tackle the class imbalance problem, as our dataset is imbalanced toward the positive sentiment class, and its aggravation by the semi-supervised learning methods, we employ data augmentation in the semi-supervised learning process in order to equalize the class distribution. Our results show that semi-supervised learning coupled with data augmentation outperforms significantly the default semi-supervised annotation process. We make the so-called TSentiment15 sentiment-annotated dataset available to the community to be used for evaluation purposes and for developing new methods.