Machine Learning to Detect Self-Reporting of Symptoms, Testing Access, and Recovery Associated With COVID-19 on Twitter: Retrospective Big Data Infoveillance Study.

Machine Learning to Detect Self-Reporting of Symptoms, Testing Access, and Recovery Associated With COVID-19 on Twitter: Retrospective Big Data Infoveillance Study.
复制标题

DOI:
10.2196/19509
复制
发表时间:
2020-06-08
影响因子:
8.5
通讯作者:
Cuomo, Raphael
Cuomo, Raphael
中科院分区:
医学3区
文献类型:
--
作者:
Mackey, Tim;Purushothaman, Vidya;Cuomo, Raphael

文献摘要

被引文献

相似文献

背景:冠状病毒病(新冠肺炎)大流行是一种全球性的紧急卫生事件,截至2020年6月初,全球已有600多万病例。鉴于它是在一个日益数字化的时代出现的,这场大流行在范围和先例上都是历史性的。重要的是,由于无法获得检测和难以测量康复等问题,人们一直担心新冠肺炎病例计数的准确性。目的:这项研究的目的是检测和表征用户生成的对话,这些对话可能与新冠肺炎相关症状、获得检测的体验以及使用无监督机器学习方法提到的疾病恢复有关。方法:从2020年3月3-20日从推特公共流媒体应用编程接口收集推文,筛选与新冠肺炎相关的通用关键字,然后进一步筛选可能与用户自我报告的新冠肺炎症状相关的术语。使用一种名为biterm主题模型(BTM)的无监督机器学习方法来分析推文,其中包含相同单词相关主题的推文组被分成主题簇,其中包括关于症状、测试和恢复的对话。结果:共收集到4,492,954条推文,包含可能与新冠肺炎症状相关的术语。在使用BTM识别相关话题聚类并删除重复推文后,我们总共识别了3465条(
BACKGROUND: The coronavirus disease (COVID-19) pandemic is a global health emergency with over 6 million cases worldwide as of the beginning of June 2020. The pandemic is historic in scope and precedent given its emergence in an increasingly digital era. Importantly, there have been concerns about the accuracy of COVID-19 case counts due to issues such as lack of access to testing and difficulty in measuring recoveries.OBJECTIVE: The aims of this study were to detect and characterize user-generated conversations that could be associated with COVID-19-related symptoms, experiences with access to testing, and mentions of disease recovery using an unsupervised machine learning approach.METHODS: Tweets were collected from the Twitter public streaming application programming interface from March 3-20, 2020, filtered for general COVID-19-related keywords and then further filtered for terms that could be related to COVID-19 symptoms as self-reported by users. Tweets were analyzed using an unsupervised machine learning approach called the biterm topic model (BTM), where groups of tweets containing the same word-related themes were separated into topic clusters that included conversations about symptoms, testing, and recovery. Tweets in these clusters were then extracted and manually annotated for content analysis and assessed for their statistical and geographic characteristics.RESULTS: A total of 4,492,954 tweets were collected that contained terms that could be related to COVID-19 symptoms. After using BTM to identify relevant topic clusters and removing duplicate tweets, we identified a total of 3465 (