Toward automated e-cigarette surveillance: Spotting e-cigarette proponents on Twitter

Toward automated e-cigarette surveillance: Spotting e-cigarette proponents on Twitter
复制标题

DOI:
10.1016/j.jbi.2016.03.006
复制
发表时间:
2016-06-01
影响因子:
4.5
通讯作者:
Sabbir, A. K. M.
Sabbir, A. K. M.
中科院分区:
医学3区
文献类型:
--
作者:
Kavuluru, Ramakanth;Sabbir, A. K. M.

文献摘要

被引文献

相似文献

背景:电子烟(电子烟或电子烟)是一种流行的新兴烟草产品。由于电子烟不会产生因吸普通香烟而产生的有毒烟草燃烧产物,它们有时被认为是吸烟的一种危害较小的替代品,也是戒烟的手段。然而,电子烟的安全性及其支持戒烟的有效性尚未确定。重要的是,联邦药品管理局(Fda)目前不对电子烟进行监管,因此其制造、营销和销售不受适用于传统香烟的规则的约束。许多制造商、倡导者和电子烟用户都在积极地在Twitter上推广电子烟。目的:我们开发了一个高精度的监督预测模型,用于自动识别Twitter上的电子烟支持者,并分析他们的推文行为随热门主题的量化变化,与其他Twitter用户(或推特用户)进行比较。方法:使用两个不同注释者的1000个独立标注的Twitter个人资料的数据集,我们利用最新推文内容和推特用户个人资料中的各种文本特征来构建预测模型,以自动识别推特支持者。我们使用一组人工精选的关键短语对来自100多万条电子CIG推文的推文进行分析,并将结果与普通推文生成的结果进行比较。结果:我们的模型识别电子CIG支持者的准确率为97%,召回率为86%,F-Score为91%,总体准确率为96%,可信区间为95%。我们发现,与构成超过90%的数据集的普通推特相反,e-cig支持者的推文子集要小得多,但推文数量是普通推特的两到五倍。支持者还不成比例地(多一到两个数量级)强调了电子烟的味道,它们的无烟和潜在的减少危害的方面,以及它们声称在戒烟中的用途。结论:鉴于FDA目前正在提出有意义的法规,我们相信我们的工作展示了信息学方法,特别是机器学习方法,在Twitter上自动监视电子烟的强大潜力。(C)2016 Elsevier Inc.保留所有权利。
Background: Electronic cigarettes (e-cigarettes or e-cigs) are a popular emerging tobacco product. Because e-cigs do not generate toxic tobacco combustion products that result from smoking regular cigarettes, they are sometimes perceived and promoted as a less harmful alternative to smoking and also as means to quit smoking. However, the safety of e-cigs and their efficacy in supporting smoking cessation is yet to be determined. Importantly, the federal drug administration (FDA) currently does not regulate e-cigs and as such their manufacturing, marketing, and sale is not subject to the rules that apply to traditional cigarettes. A number of manufacturers, advocates, and e-cig users are actively promoting e-cigs on Twitter.Objective: We develop a high accuracy supervised predictive model to automatically identify e-cig "proponents" on Twitter and analyze the quantitative variation of their tweeting behavior along popular themes when compared with other Twitter users (or tweeters).Methods: Using a dataset of 1000 independently annotated Twitter profiles by two different annotators, we employed a variety of textual features from latest tweet content and tweeter profile biography to build predictive models to automatically identify proponent tweeters. We used a set of manually curated key phrases to analyze e-cig proponent tweets from a corpus of over one million e-cig tweets along well known e-cig themes and compared the results with those generated by regular tweeters.Results: Our model identifies e-cig proponents with 97% precision, 86% recall, 91% F-score, and 96% overall accuracy, with tight 95% confidence intervals. We find that as opposed to regular tweeters that form over 90% of the dataset, e-cig proponents are a much smaller subset but tweet two to five times more than regular tweeters. Proponents also disproportionately (one to two orders of magnitude more) highlight e-cig flavors, their smoke-free and potential harm reduction aspects, and their claimed use in smoking cessation.Conclusions: Given FDA is currently in the process of proposing meaningful regulation, we believe our work demonstrates the strong potential of informatics approaches, specifically machine learning, for automated e-cig surveillance on Twitter. (C) 2016 Elsevier Inc. All rights reserved.