Using twitter to examine smoking behavior and perceptions of emerging tobacco products.

Using twitter to examine smoking behavior and perceptions of emerging tobacco products.
复制标题

DOI:
10.2196/jmir.2534
复制
发表时间:
2013-08-29
影响因子:
7.4
通讯作者:
Conway M
Conway M
中科院分区:
医学2区
文献类型:
--
作者:
Myslín M;Zhu SH;Chapman W;Conway M

文献摘要

参考文献

被引文献

相似文献

Twitter等社交媒体平台正迅速成为公共卫生监测应用的关键资源,但人们对Twitter用户对烟草的信息和情绪水平知之甚少,特别是关于水烟和电子烟带来的新烟草控制挑战。开发与烟草相关的Twitter帖子的内容和情感分析,并构建机器学习分类器来检测与烟草相关的帖子和对烟草的情感,特别关注水烟和电子烟等新兴产品。从2011年12月到2012年7月,我们每隔15天收集了7362条与烟草相关的Twitter帖子。每条推文都使用三轴方案进行手动分类,捕获流派,主题和情绪。使用收集的数据,机器学习分类器被训练来检测烟草相关与不相关的推文以及积极与消极的情绪,使用朴素贝叶斯,k最近邻和支持向量机(SVM)算法。最后,计算每个类别之间的phi列联系数以发现涌现模式。最流行的类型是第一手和第二手的经验和意见,最常见的主题是水烟,停止和快乐。总体而言,对烟草的态度更积极(1939/4215,46%的推文),而不是负面(1349/4215,32%)或中性的推文提到它,甚至排除了9%的推文归类为营销。三个独立的指标融合在一起,以支持一个紧急的区别,一方面,水烟和电子烟对应于积极的情绪,另一方面,传统的烟草产品和更一般的参考对应于消极的情绪。这些指标包括注释方案中类别之间的相关性(phihookah-positive=0.39; phie-cigs-positive=0.19);搜索关键词和情感之间的相关性(χ2 4=414.50,P<0.001,Cramer's V=0.36),以及在研究的机器学习组件中通过对数比值比排名的积极和消极情感的最具鉴别力的单字特征。在自动分类任务中,使用相对少量的unigram特征(500)的SVM在区分烟草相关和不相关的推文方面取得了最佳性能(F得分=0.85)。通过Twitter进行烟草监测的新见解通过积极情绪的高度流行得到了证明。这种积极的情绪以复杂的方式与社会形象、个人经历以及最近流行的产品(如水烟和电子烟)相关。这些产品与其健康影响之间的一些明显的感知脱节表明了烟草控制教育的机会。最后,烟草相关帖子的机器分类显示出比严格基于关键词的方法更有希望的优势,提高了Twitter数据的信噪比,并为自动烟草监控应用铺平了道路。
Social media platforms such as Twitter are rapidly becoming key resources for public health surveillance applications, yet little is known about Twitter users’ levels of informedness and sentiment toward tobacco, especially with regard to the emerging tobacco control challenges posed by hookah and electronic cigarettes. To develop a content and sentiment analysis of tobacco-related Twitter posts and build machine learning classifiers to detect tobacco-relevant posts and sentiment towards tobacco, with a particular focus on new and emerging products like hookah and electronic cigarettes. We collected 7362 tobacco-related Twitter posts at 15-day intervals from December 2011 to July 2012. Each tweet was manually classified using a triaxial scheme, capturing genre, theme, and sentiment. Using the collected data, machine-learning classifiers were trained to detect tobacco-related vs irrelevant tweets as well as positive vs negative sentiment, using Naïve Bayes, k-nearest neighbors, and Support Vector Machine (SVM) algorithms. Finally, phi contingency coefficients were computed between each of the categories to discover emergent patterns. The most prevalent genres were first- and second-hand experience and opinion, and the most frequent themes were hookah, cessation, and pleasure. Sentiment toward tobacco was overall more positive (1939/4215, 46% of tweets) than negative (1349/4215, 32%) or neutral among tweets mentioning it, even excluding the 9% of tweets categorized as marketing. Three separate metrics converged to support an emergent distinction between, on one hand, hookah and electronic cigarettes corresponding to positive sentiment, and on the other hand, traditional tobacco products and more general references corresponding to negative sentiment. These metrics included correlations between categories in the annotation scheme (phihookah-positive=0.39; phie-cigs-positive=0.19); correlations between search keywords and sentiment (χ2 4=414.50, P<.001, Cramer’s V=0.36), and the most discriminating unigram features for positive and negative sentiment ranked by log odds ratio in the machine learning component of the study. In the automated classification tasks, SVMs using a relatively small number of unigram features (500) achieved best performance in discriminating tobacco-related from unrelated tweets (F score=0.85). Novel insights available through Twitter for tobacco surveillance are attested through the high prevalence of positive sentiment. This positive sentiment is correlated in complex ways with social image, personal experience, and recently popular products such as hookah and electronic cigarettes. Several apparent perceptual disconnects between these products and their health effects suggest opportunities for tobacco control education. Finally, machine classification of tobacco-related posts shows a promising edge over strictly keyword-based approaches, yielding an improved signal-to-noise ratio in Twitter data and paving the way for automated tobacco surveillance applications.
DOI: 10.1371/journal.pone.0014118
发表时间: 2010-11-29
期刊: PloS one
影响因子: 3.7
作者:
Chew C;Eysenbach G
通讯作者: Eysenbach G
DOI: 10.1093/ntr/ntq212
发表时间: 2011-02-01
影响因子: 4.7
作者:
Cobb, Caroline O.;Shihadeh, Alan;Eissenberg, Thomas
通讯作者: Eissenberg, Thomas
DOI: 10.1177/0022034511415273
发表时间: 2011-09-01
影响因子: 7.6
作者:
Heaivilin, N.;Gerbert, B.;Gibbs, J. L.
通讯作者: Gibbs, J. L.
DOI: 10.1111/j.1742-1241.2011.02751.x
发表时间: 2011-10-01
影响因子: 2.6
作者:
Foulds, J.;Veldheer, S.;Berg, A.
通讯作者: Berg, A.
DOI: 10.1136/tc.2010.042507
发表时间: 2012-07-01
期刊: TOBACCO CONTROL
影响因子: 5.2
作者:
Prochaska, Judith J.;Pechmann, Cornelia;Leonhardt, James M.
通讯作者: Leonhardt, James M.