Using Deep Learning to Identify Linguistic Features that Facilitate or Inhibit the Propagation of Anti- and Pro-Vaccine Content on Social Media.

Using Deep Learning to Identify Linguistic Features that Facilitate or Inhibit the Propagation of Anti- and Pro-Vaccine Content on Social Media.
复制标题

使用深度学习来识别促进或抑制社交媒体上反对和支持疫苗内容传播的语言特征。

DOI:
10.1109/icdh55609.2022.00025
复制
发表时间:
2022
期刊:
2022 IEEE International Conference on Digital Health (IEEE ICDH 2022) : proceedings : hybrid conference, Barcelona, Spain, 11-15 July 2022. International Conference on Digital Health (2022 : Barcelona, Spain; Online)
影响因子:
--
通讯作者:
Tan,Pang-Ning
Tan,Pang-Ning
中科院分区:
--
文献类型:
--
作者:
Argyris,YoungAnna;Zhang,Nan;Bashyal,Bidhan;Tan,Pang-Ning

文献摘要

相似文献

反疫苗的内容通过社交媒体迅速传播,助长了疫苗的犹豫不决,而支持疫苗的内容并没有复制对手的成功。尽管反对和支持疫苗的帖子在传播方面存在差异,但促进或抑制疫苗相关内容传播的语言特征仍然鲜为人知。此外,大多数以前的机器学习算法将社交媒体帖子分为两类(例如,错误信息或非错误信息),并且很少基于对疫苗的不同观点(例如,反疫苗、支持疫苗和中性)来处理更高级别的分类任务。我们的目标是(1)确定促进和抑制疫苗相关内容传播的语言特征集,以及(2)比较反对疫苗、支持疫苗和中立的推文是否比其他推文包含更频繁的一组内容。为了实现这些目标,我们在2021年11月15日至12月15日期间收集了大量社交媒体帖子(超过1.2亿条推文),恰逢奥米克隆变种的激增。使用微调的BERT分类器开发了一个两阶段框架,对二元和三元分类的准确率分别超过99%和80%。最后,使用语言查询字数统计文本分析工具统计每条分类推文中的语言特征。我们的回归结果显示,反对疫苗的推文被传播(即,转发),而支持疫苗的推文获得被动支持(即,受欢迎)。我们的结果还产生了两组语言特征,作为疫苗相关推文传播的促进者和抑制者。最后,我们的回归结果显示,反对疫苗的推文倾向于使用促进剂,而支持疫苗的推文则使用抑制剂。这项研究的这些发现和算法将有助于公共卫生官员努力消除疫苗错误信息,从而促进在大流行和流行病期间提供预防措施。
Anti-vaccine content is rapidly propagated via social media, fostering vaccine hesitancy, while pro-vaccine content has not replicated the opponent's successes. Despite this dis-parity in the dissemination of anti- and pro-vaccine posts, linguistic features that facilitate or inhibit the propagation of vaccine-related content remain less known. Moreover, most prior machine-learning algorithms classified social-media posts into binary categories (e.g., misinformation or not) and have rarely tackled a higher-order classification task based on divergent perspectives about vaccines (e.g., anti-vaccine, pro-vaccine, and neutral). Our objectives are (1) to identify sets of linguistic features that facilitate and inhibit the propagation of vaccine-related content and (2) to compare whether anti-vaccine, pro-vaccine, and neutral tweets contain either set more frequently than the others. To achieve these goals, we collected a large set of social media posts (over 120 million tweets) between Nov. 15 and Dec. 15, 2021, coinciding with the Omicron variant surge. A two-stage framework was developed using a fine-tuned BERT classifier, demonstrating over 99 and 80 percent accuracy for binary and ternary classification. Finally, the Linguistic Inquiry Word Count text analysis tool was used to count linguistic features in each classified tweet. Our regression results show that anti-vaccine tweets are propagated (i.e., retweeted), while pro-vaccine tweets garner passive endorsements (i.e., favorited). Our results also yielded the two sets of linguistic features as facilitators and inhibitors of the propagation of vaccine-related tweets. Finally, our regression results show that anti-vaccine tweets tend to use the facilitators, while pro-vaccine counterparts employ the inhibitors. These findings and algorithms from this study will aid public health officials' efforts to counteract vaccine misinformation, thereby facilitating the delivery of preventive measures during pandemics and epidemics.