Machine Learning Maps Research Needs in COVID-19 Literature

Machine Learning Maps Research Needs in COVID-19 Literature
复制标题

DOI:
10.1016/j.patter.2020.100123
复制
发表时间:
2020-12-11
期刊:
影响因子:
6.5
通讯作者:
Majumder, Maimuna
Majumder, Maimuna
中科院分区:
其他
文献类型:
--
作者:
Doanvo, Anhvinh;Qian, Xiaolu;Majumder, Maimuna

文献摘要

被引文献

相似文献

截至 2020 年 8 月,已制作了数千份有关 COVID-19(2019 年冠状病毒病)的出版物。手动评估其范围是一项艰巨的任务,而通过元数据分析(例如关键字)的捷径假设研究已被正确标记。然而,机器学习方法可以快速调查出版物摘要的实际文本,以识别 COVID-19 与其他冠状病毒之间的研究重叠、研究热点和值得探索的领域。我们提出了一个快速、可扩展且可重用的框架来解析新疾病文献。当应用于 COVID-19 开放研究数据集时,降维表明迄今为止的 COVID-19 研究主要是基于临床、建模或现场的,这与针对其他(非 COVID-19)冠状病毒疾病的大量实验室驱动的研究形成鲜明对比。此外,主题模型表明,COVID-19 出版物的重点是公共卫生、疫情报告、临床护理和冠状病毒检测,而不是关注基础微生物学(包括发病机制和传播)的出版物数量更为有限。
As of August 2020, thousands of COVID-19 (coronavirus disease 2019) publications have been produced. Manual assessment of their scope is an overwhelming task, and shortcuts through metadata analysis (e.g., keywords) assume that studies are properly tagged. However, machine learning approaches can rapidly survey the actual text of publication abstracts to identify research overlap between COVID-19 and other coronaviruses, research hotspots, and areas warranting exploration. We propose a fast, scalable, and reusable framework to parse novel disease literature. When applied to the COVID-19 Open Research Dataset, dimensionality reduction suggests that COVID-19 studies to date are primarily clinical, modeling, or field based, in contrast to the vast quantity of laboratory-driven research for other (non-COVID-19) coronavirus diseases. Furthermore, topic modeling indicates that COVID-19 publications have focused on public health, outbreak reporting, clinical care, and testing for coronaviruses, as opposed to the more limited number focused on basic microbiology, including pathogenesis and transmission.