Identifying Electronic Nicotine Delivery System Brands and Flavors on Instagram: Natural Language Processing Analysis.

Identifying Electronic Nicotine Delivery System Brands and Flavors on Instagram: Natural Language Processing Analysis.
复制标题

DOI:
10.2196/30257
复制
发表时间:
2022-01-18
影响因子:
7.4
通讯作者:
Kim A
Kim A
中科院分区:
医学2区
文献类型:
--
作者:
Chew R;Wenger M;Guillory J;Nonnemaker J;Kim A

文献摘要

参考文献

被引文献

相似文献

电子尼古丁输送系统(ENDS)品牌,如JUUL,将社交媒体作为其营销策略的关键组成部分,这导致了2015年至2018年的大规模销售增长。在此期间,ENDS的使用在青少年和年轻人中迅速增加,调味产品在这些群体中特别受欢迎。我们研究的目的是开发一个命名实体识别(NER)模型,以从Instagram帖子文本中识别潜在的新兴电子烟品牌和口味。NER是一种自然语言处理任务,用于基于实体和周围单词的特征来识别文本中特定类型的单词(实体)。NER模型在2272个Instagram帖子的标记数据集上进行训练,这些帖子针对ENDS品牌和口味进行编码。我们比较了三种类型的NER模型--条件随机场、残差卷积神经网络和来自变压器(FTDB)网络的微调蒸馏双向编码器表示--以识别Instagram帖子中的品牌和口味,这些帖子具有精确度、召回率和F1分数的关键模型结果。我们使用尼尔森扫描仪销售和维基百科的数据创建基准词典,以确定在我们的样本中,Instagram帖子中是否提到了ENDS品牌和风味列表中的品牌。为了防止过度拟合,我们进行了5折交叉验证,并报告了跨折的模型验证指标的平均值和SD。对于品牌,残差卷积神经网络表现出最高的平均精度(0.797,SD 0.084),FTDB表现出最高的平均召回率(0.869,SD 0.103)。对于风味,FTDB表现出最高的平均精确度(0.860,SD 0.055)和召回(0.801,SD 0.091)。所有NER模型在平均精确度、召回率和F1方面都优于基准品牌和风味词典查找。比较基准品牌列表,较大的维基百科列表在准确率和召回率方面都优于尼尔森列表。我们的研究结果表明,NER模型正确识别了Instagram帖子中的ENDS品牌和口味,其识别率与已发表文献中的其他品牌和口味相比具有竞争力或更好。在人工注释过程中识别的品牌与尼尔森扫描仪数据中的品牌几乎没有重叠,这表明NER模型可能会捕捉到销售和分销有限的新兴品牌。NER模型解决了人工品牌识别的挑战,并可用于支持未来的信息学和信息传播学研究。社交媒体上识别的品牌应与尼尔森和其他数据源进行交叉验证,以区分已建立的新兴品牌和销售和分销有限的品牌。
Electronic nicotine delivery system (ENDS) brands, such as JUUL, used social media as a key component of their marketing strategy, which led to massive sales growth from 2015 to 2018. During this time, ENDS use rapidly increased among youths and young adults, with flavored products being particularly popular among these groups. The aim of our study is to develop a named entity recognition (NER) model to identify potential emerging vaping brands and flavors from Instagram post text. NER is a natural language processing task for identifying specific types of words (entities) in text based on the characteristics of the entity and surrounding words. NER models were trained on a labeled data set of 2272 Instagram posts coded for ENDS brands and flavors. We compared three types of NER models—conditional random fields, a residual convolutional neural network, and a fine-tuned distilled bidirectional encoder representations from transformers (FTDB) network—to identify brands and flavors in Instagram posts with key model outcomes of precision, recall, and F1 scores. We used data from Nielsen scanner sales and Wikipedia to create benchmark dictionaries to determine whether brands from established ENDS brand and flavor lists were mentioned in the Instagram posts in our sample. To prevent overfitting, we performed 5-fold cross-validation and reported the mean and SD of the model validation metrics across the folds. For brands, the residual convolutional neural network exhibited the highest mean precision (0.797, SD 0.084), and the FTDB exhibited the highest mean recall (0.869, SD 0.103). For flavors, the FTDB exhibited both the highest mean precision (0.860, SD 0.055) and recall (0.801, SD 0.091). All NER models outperformed the benchmark brand and flavor dictionary look-ups on mean precision, recall, and F1. Comparing between the benchmark brand lists, the larger Wikipedia list outperformed the Nielsen list in both precision and recall. Our findings suggest that NER models correctly identified ENDS brands and flavors in Instagram posts at rates competitive with, or better than, others in the published literature. Brands identified during manual annotation showed little overlap with those in Nielsen scanner data, suggesting that NER models may capture emerging brands with limited sales and distribution. NER models address the challenges of manual brand identification and can be used to support future infodemiology and infoveillance studies. Brands identified on social media should be cross-validated with Nielsen and other data sources to differentiate emerging brands that have become established from those with limited sales and distribution.
DOI: 10.1073/pnas.1907367117
发表时间: 2020-12-01
影响因子: 11.1
作者:
Manning, Christopher D.;Clark, Kevin;Levy, Omer
通讯作者: Levy, Omer
DOI: 10.15585/mmwr.mm6745a5
发表时间: 2018-11-16
期刊: MMWR. Morbidity and mortality weekly report
影响因子: --
作者:
Cullen KA;Ambrose BK;Gentzke AS;Apelberg BJ;Jamal A;King BA
通讯作者: King BA
DOI: 10.15585/mmwr.mm6946a4
发表时间: 2020-11-20
期刊: MMWR. Morbidity and mortality weekly report
影响因子: --
作者:
Cornelius ME;Wang TW;Jamal A;Loretan CG;Neff LJ
通讯作者: Neff LJ
DOI: 10.1377/hlthaff.24.6.1601
发表时间: 2005-11-01
期刊: HEALTH AFFAIRS
影响因子: 9.7
作者:
Carpenter, CM;Wayne, GF;Connolly, GN
通讯作者: Connolly, GN
DOI: 10.2196/jmir.8550
发表时间: 2018-03-01
影响因子: 7.4
作者:
Hsu, Greta;Sun, Jessica Y.;Zhu, Shu-Hong
通讯作者: Zhu, Shu-Hong