Big Data, Natural Language Processing, and Deep Learning to Detect and Characterize Illicit COVID-19 Product Sales: Infoveillance Study on Twitter and Instagram

Big Data, Natural Language Processing, and Deep Learning to Detect and Characterize Illicit COVID-19 Product Sales: Infoveillance Study on Twitter and Instagram
复制标题

DOI:
10.2196/20794
复制
发表时间:
2020-07-01
影响因子:
8.5
通讯作者:
Liang, Bryan
Liang, Bryan
中科院分区:
医学3区
文献类型:
--
作者:
Mackey, Tim Ken;Li, Jiawei;Liang, Bryan

文献摘要

被引文献

相似文献

背景:冠状病毒病(COVID-19)大流行可能是上个世纪最大的全球健康挑战。伴随这场大流行的是一种平行的“信息流行病”,包括在线营销和销售未经批准的、非法的和假冒的COVID-19保健产品,包括检测试剂盒、治疗方法和其他可疑的“治疗方法”。“基于互联网的技术日益普及,包括现在拥有数十亿全球用户的流行社交媒体平台,使这种内容的扩散成为可能。目的:本研究旨在收集、分析、识别和报告来自Twitter和Instagram的疑似假冒、伪造和未经批准的COVID-19相关医疗保健产品。方法:这项研究分两个阶段进行,首先是收集与COVID-19相关的Twitter和Instagram帖子,使用Instagram上的网络抓取和过滤公共流媒体Twitter应用程序编程接口中与可疑营销和销售COVID-19相关的关键字。19个产品。第二阶段涉及使用自然语言处理(NLP)和深度学习进行数据分析,以识别潜在的卖家,然后手动注释感兴趣的特征。我们还在定制的数据仪表板上可视化非法销售帖子,以实现公共卫生情报。结果:我们收集了6,029,323条推文和204,597条Instagram帖子,过滤了3月至4月Twitter和2月至5月Instagram与可疑营销和销售COVID-19保健产品相关的术语。在应用我们的NLP和深度学习方法后,我们发现了1271条推文和596条Instagram帖子与COVID-19相关产品的可疑销售有关。一般来说,产品的推出分为两波,第一波是可疑的免疫增强治疗,第二波是可疑的检测试剂盒。我们还检测到少量尚未获批用于COVID-19治疗的药物。发现的其他主要主题包括以不同语言提供的产品、各种产品可信度声明、完全未经证实的产品、未经批准的测试方式以及不同的付款和卖家联系方式。这项研究的结果通过描述什么类型的健康产品,销售索赔,在疫情早期阶段,两个流行的社交媒体平台上活跃着不同类型的卖家。随着疫情的发展以及更多人寻求COVID-19检测和治疗,这一网络犯罪挑战可能会持续下去。这种数据智能可以帮助公共卫生机构、监管机构、合法制造商和技术平台更好地删除和防止这些内容伤害公众。
Background: The coronavirus disease (COVID-19) pandemic is perhaps the greatest global health challenge of the last century. Accompanying this pandemic is a parallel "infodemic," including the online marketing and sale of unapproved, illegal, and counterfeit COVID-19 health products including testing kits, treatments, and other questionable "cures." Enabling the proliferation of this content is the growing ubiquity of internet-based technologies, including popular social media platforms that now have billions of global users.Objective: This study aims to collect, analyze, identify, and enable reporting of suspected fake, counterfeit, and unapproved COVID-19-related health care products from Twitter and Instagram.Methods: This study is conducted in two phases beginning with the collection of COVID-19-related Twitter and Instagram posts using a combination of web scraping on Instagram and filtering the public streaming Twitter application programming interface for keywords associated with suspect marketing and sale of COVID-19 products. The second phase involved data analysis using natural language processing (NLP) and deep learning to identify potential sellers that were then manually annotated for characteristics of interest. We also visualized illegal selling posts on a customized data dashboard to enable public health intelligence.Results: We collected a total of 6,029,323 tweets and 204,597 Instagram posts filtered for terms associated with suspect marketing and sale of COVID-19 health products from March to April for Twitter and February to May for Instagram. After applying our NLP and deep learning approaches, we identified 1271 tweets and 596 Instagram posts associated with questionable sales of COVID-19-related products. Generally, product introduction came in two waves, with the first consisting of questionable immunity-boosting treatments and a second involving suspect testing kits. We also detected a low volume of pharmaceuticals that have not been approved for COVID-19 treatment. Other major themes detected included products offered in different languages, various claims of product credibility, completely unsubstantiated products, unapproved testing modalities, and different payment and seller contact methods.Conclusions: Results from this study provide initial insight into one front of the "infodemic" fight against COVID-19 by characterizing what types of health products, selling claims, and types of sellers were active on two popular social media platforms at earlier stages of the pandemic. This cybercrime challenge is likely to continue as the pandemic progresses and more people seek access to COVID-19 testing and treatment. This data intelligence can help public health agencies, regulatory authorities, legitimate manufacturers, and technology platforms better remove and prevent this content from harming the public.