Using Natural Language Processing to Explore "Dry January" Posts on Twitter: Longitudinal Infodemiology Study.

Using Natural Language Processing to Explore "Dry January" Posts on Twitter: Longitudinal Infodemiology Study.
复制标题

DOI:
10.2196/40160
复制
发表时间:
2022-11-18
影响因子:
7.4
通讯作者:
Massey, Philip M.
Massey, Philip M.
中科院分区:
医学2区
文献类型:
--
作者:
Russell, Alex M.;Valdez, Danny;Chiang, Shawn C.;Montemayor, Ben N.;Barry, Adam E.;Lin, Hsien-Chang;Massey, Philip M.

文献摘要

参考文献

被引文献

相似文献

“一月戒酒”是一项临时禁酒运动,鼓励人们通过在一月份暂时戒酒来反思他们与酒精的关系。尽管“一月戒酒”已经成为一种全球现象,但对“一月戒酒”参与者经历的调查却很有限。要想深入了解个人“一月戒酒”相关经历,一种方法是利用大规模的社交媒体数据(如Twitter聊天)来探索和描述有关“一月戒酒”的公共话语。我们试图回答以下问题:(1)关于“一月戒酒”的推文语料库中存在哪些主题,以及在多年推文(2020-2022)中用于讨论“一月戒酒”的语言是否存在一致性?(2) 2019冠状病毒病大流行爆发后,“2021年1月干”推文中是否出现了独特的主题或模式?(3)推文组成(即情绪、人工撰写与机器人撰写)与Dry January推文的参与度有何关联?我们将自然语言处理技术应用于大量推文样本(n=222,917),其中包含12月15日至2月15日在三个独立的参与年份(2020-2022)中发布的术语“dryjanuary”或“dryjanuary”。使用术语频率逆文档频率、k-均值聚类和主成分分析进行数据可视化,以确定每年的最佳聚类数量。一旦数据可视化,我们运行解释模型以提供年内(或集群内)比较。使用潜狄利克雷分配主题模型来检查每个给定年份的每个簇内的内容。使用价感知词典和情感推理器情感分析来检查每年每个聚类的影响。使用Botometer自动帐户检查来确定每年每个集群的平均bot分数。最后,为了评估Dry January内容的用户参与度,我们取了每个集群的平均点赞数和转发数,并与其他感兴趣的结果变量进行了相关性分析。我们每年观察到几个类似的主题(例如,1月干燥资源,1月干燥健康益处,与1月干燥进展相关的更新),表明1月干燥内容随着时间的推移相对一致。尽管多年推文的主题存在重叠,但在2021年的推文语料库中发现了与2019冠状病毒病全球大流行期间个人饮酒经历相关的独特主题。此外,推文的构成与参与度有关,包括每条推文的点赞、转发和引用数量。与人类创作的集群相比,机器人主导的集群有更少的点赞、转发或引用推文。研究结果强调了使用大规模社交媒体(如Twitter上的讨论)来研究减少饮酒的尝试,并监测正在考虑、准备或积极尝试戒烟或减少饮酒的人的持续动态需求的效用。
Dry January, a temporary alcohol abstinence campaign, encourages individuals to reflect on their relationship with alcohol by temporarily abstaining from consumption during the month of January. Though Dry January has become a global phenomenon, there has been limited investigation into Dry January participants’ experiences. One means through which to gain insights into individuals’ Dry January-related experiences is by leveraging large-scale social media data (eg, Twitter chatter) to explore and characterize public discourse concerning Dry January. We sought to answer the following questions: (1) What themes are present within a corpus of tweets about Dry January, and is there consistency in the language used to discuss Dry January across multiple years of tweets (2020-2022)? (2) Do unique themes or patterns emerge in Dry January 2021 tweets after the onset of the COVID-19 pandemic? and (3) What is the association with tweet composition (ie, sentiment and human-authored vs bot-authored) and engagement with Dry January tweets? We applied natural language processing techniques to a large sample of tweets (n=222,917) containing the term “dry january” or “dryjanuary” posted from December 15 to February 15 across three separate years of participation (2020-2022). Term frequency inverse document frequency, k-means clustering, and principal component analysis were used for data visualization to identify the optimal number of clusters per year. Once data were visualized, we ran interpretation models to afford within-year (or within-cluster) comparisons. Latent Dirichlet allocation topic modeling was used to examine content within each cluster per given year. Valence Aware Dictionary and Sentiment Reasoner sentiment analysis was used to examine affect per cluster per year. The Botometer automated account check was used to determine average bot score per cluster per year. Last, to assess user engagement with Dry January content, we took the average number of likes and retweets per cluster and ran correlations with other outcome variables of interest. We observed several similar topics per year (eg, Dry January resources, Dry January health benefits, updates related to Dry January progress), suggesting relative consistency in Dry January content over time. Although there was overlap in themes across multiple years of tweets, unique themes related to individuals’ experiences with alcohol during the midst of the COVID-19 global pandemic were detected in the corpus of tweets from 2021. Also, tweet composition was associated with engagement, including number of likes, retweets, and quote-tweets per post. Bot-dominant clusters had fewer likes, retweets, or quote tweets compared with human-authored clusters. The findings underscore the utility for using large-scale social media, such as discussions on Twitter, to study drinking reduction attempts and to monitor the ongoing dynamic needs of persons contemplating, preparing for, or actively pursuing attempts to quit or cut down on their drinking.
DOI: 10.1093/eurpub/ckx124
发表时间: 2017-10-01
影响因子: 4.4
作者:
de Visser, Richard O.;Robinson, Emily;Walmsley, Matthew
通讯作者: Walmsley, Matthew
DOI: 10.1080/08870446.2020.1743840
发表时间: 2020-03-26
影响因子: 3.3
作者:
de Visser, Richard O.;Nicholls, James
通讯作者: Nicholls, James
DOI: 10.1016/j.drugalcdep.2021.108672
发表时间: 2021-05-01
影响因子: 4.2
作者:
Bunting AM;Frank D;Arshonsky J;Bragg MA;Friedman SR;Krawczyk N
通讯作者: Krawczyk N
DOI: 10.1080/09638237.2019.1677878
发表时间: 2019-11-06
影响因子: 3.3
作者:
Budenz, Alexandra;Klassen, Ann;Massey, Philip
通讯作者: Massey, Philip
DOI: 10.2196/jmir.3622
发表时间: 2014-10-16
影响因子: 7.4
作者:
Harris JK;Moreland-Russell S;Choucair B;Mansour R;Staub M;Simmons K
通讯作者: Simmons K