Identifying Sentiment of Hookah-Related Posts on Twitter.

Identifying Sentiment of Hookah-Related Posts on Twitter.
复制标题

DOI:
10.2196/publichealth.8133
复制
发表时间:
2017-10-18
影响因子:
8.5
通讯作者:
Unger JB
Unger JB
中科院分区:
医学3区
文献类型:
--
作者:
Allem JP;Ramanujam J;Lerman K;Chu KH;Boley Cruz T;Unger JB

文献摘要

被引文献

相似文献

水烟(或水管)在美国和其他地方越来越受欢迎,对公共卫生产生了影响,因为它具有与可燃香烟相似的健康风险。虽然水烟的使用迅速普及,但社交媒体数据(Twitter,Instagram)可用于捕获和描述个人使用,感知,讨论和销售这种烟草产品的社会和环境背景。这些数据可以让人们有机地报告他们对烟草产品的看法,如研究人员未准备的水烟,没有仪器偏见,成本低。这项研究描述了Twitter上与水烟相关的帖子的情绪,并描述了在试图理解态度时消除Twitter数据的重要性。从2015年3月24日到2016年12月2日,在Twitter上收集了与Hookah相关的帖子(N= 986,320)。机器学习模型被用来描述20种不同情绪的情绪,并对数据进行去偏置,以便Twitter帖子反映合法人类用户的情绪,而不是社交机器人或营销型账户的情绪,这些账户可能会提供过于积极或过于消极的水烟情绪。从分析样本中,352,116条推文(59.50%)被归类为积极的,177,537条(30.00%)被归类为消极的,62,139条(10.50%)中性的。在所有积极的推文中,218,312(62.00%)被归类为高度积极的情绪(例如,积极,警觉,兴奋,兴高采烈,快乐和愉快),而133,804(38.00%)积极的推文被归类为消极的积极情绪(例如,满足,宁静,平静,放松和柔和)。在所有负面推文中,95,870(54.00%)被归类为压抑的负面情绪(例如,悲伤,不快乐,沮丧和无聊),而其余81,667(46.00%)负面推文被归类为高度负面情绪(例如,紧张,紧张,压力,不安和不愉快)。当比较有社交机器人的推文语料库和没有社交机器人的推文语料库时,情绪发生了巨大变化。例如,任何一条反映喜悦的推文的概率是61.30%,来自无偏见(或无机器人)的推文语料库。相比之下,在有偏见的语料库中,任何一条推特反映喜悦的概率为16.40%。社交媒体数据使研究人员能够通过倾听人们用自己的话说什么来了解公众的情绪和态度。负责风险沟通的烟草控制程序员可以考虑针对在Twitter上发布关于水烟的积极信息的个人,或设计放大负面情绪的信息。Twitter上传达对水烟的积极情绪的帖子可能会促进水烟使用的正常化,并且是未来研究的一个领域。这项研究的结果表明,在试图从Twitter数据中理解态度时,去偏见数据的重要性。
The increasing popularity of hookah (or waterpipe) use in the United States and elsewhere has consequences for public health because it has similar health risks to that of combustible cigarettes. While hookah use rapidly increases in popularity, social media data (Twitter, Instagram) can be used to capture and describe the social and environmental contexts in which individuals use, perceive, discuss, and are marketed this tobacco product. These data may allow people to organically report on their sentiment toward tobacco products like hookah unprimed by a researcher, without instrument bias, and at low costs. This study describes the sentiment of hookah-related posts on Twitter and describes the importance of debiasing Twitter data when attempting to understand attitudes. Hookah-related posts on Twitter (N=986,320) were collected from March 24, 2015, to December 2, 2016. Machine learning models were used to describe sentiment on 20 different emotions and to debias the data so that Twitter posts reflected sentiment of legitimate human users and not of social bots or marketing-oriented accounts that would possibly provide overly positive or overly negative sentiment of hookah. From the analytical sample, 352,116 tweets (59.50%) were classified as positive while 177,537 (30.00%) were classified as negative, and 62,139 (10.50%) neutral. Among all positive tweets, 218,312 (62.00%) were classified as highly positive emotions (eg, active, alert, excited, elated, happy, and pleasant), while 133,804 (38.00%) positive tweets were classified as passive positive emotions (eg, contented, serene, calm, relaxed, and subdued). Among all negative tweets, 95,870 (54.00%) were classified as subdued negative emotions (eg, sad, unhappy, depressed, and bored) while the remaining 81,667 (46.00%) negative tweets were classified as highly negative emotions (eg, tense, nervous, stressed, upset, and unpleasant). Sentiment changed drastically when comparing a corpus of tweets with social bots to one without. For example, the probability of any one tweet reflecting joy was 61.30% from the debiased (or bot free) corpus of tweets. In contrast, the probability of any one tweet reflecting joy was 16.40% from the biased corpus. Social media data provide researchers the ability to understand public sentiment and attitudes by listening to what people are saying in their own words. Tobacco control programmers in charge of risk communication may consider targeting individuals posting positive messages about hookah on Twitter or designing messages that amplify the negative sentiments. Posts on Twitter communicating positive sentiment toward hookah could add to the normalization of hookah use and is an area of future research. Findings from this study demonstrated the importance of debiasing data when attempting to understand attitudes from Twitter data.