Use of Machine Learning and Artificial Intelligence to predict SARS-CoV-2 infection from Full Blood Counts in a population

Use of Machine Learning and Artificial Intelligence to predict SARS-CoV-2 infection from Full Blood Counts in a population
复制标题

DOI:
10.1016/j.intimp.2020.106705
复制
发表时间:
2020-09-01
影响因子:
5.6
通讯作者:
Mackenzie, Louise S.
Mackenzie, Louise S.
中科院分区:
医学2区
文献类型:
--
作者:
Banerjee, Abhirup;Ray, Surajit;Mackenzie, Louise S.

文献摘要

被引文献

相似文献

自2019年12月以来,新型冠状病毒SARS-CoV-2被确定为导致COVID-19大流行的原因。早期症状与普通感冒和流感等其他常见疾病重叠,因此早期筛查和诊断是卫生从业人员的关键目标。该研究的目的是利用机器学习(ML)、人工神经网络(ANN)和简单的统计检验,在不了解个人症状或病史的情况下,从全血细胞计数中识别出SARS-CoV-2阳性患者。分析和培训中包含的数据集包含了在巴西圣保罗以色列阿尔伯特·爱因斯坦医院就诊的患者的匿名全血细胞计数结果,这些患者在访问医院期间收集了样本进行SARS-CoV-2 rt-PCR检测。患者数据由医院匿名化,临床数据标准化,平均值为零,单位标准差。这些数据被公开,目的是让研究人员能够开发出方法,使医院能够快速预测和潜在地识别SARS-CoV-2阳性患者。我们发现,在全血细胞计数随机森林、浅学习和灵活的人工神经网络模型下,在常规病房人群(AUC = 94-95%)和未住院或社区人群(AUC = 80-86%)之间预测SARS-CoV-2患者的准确率很高。这里,AUC是接收器工作特性曲线下的面积,是模型性能的度量。此外,使用4个血细胞计数的简单线性组合可以使社区内患者的AUC达到85%。SARS-CoV-2阳性患者不同血液参数的归一化数据显示血小板、白细胞、嗜酸性粒细胞、嗜碱性粒细胞和淋巴细胞减少,单核细胞增加。SARS-CoV-2阳性患者表现出特征性的免疫反应谱模式,并且通过简单和快速血液检查检测到的全血细胞计数测量的不同参数发生变化。虽然已知感染早期的症状与其他常见情况重叠,但可以分析全血细胞计数参数,以便在比目前针对SARS-CoV-2的rt-PCR检测所允许的更早阶段区分病毒类型。这种新方法有潜力极大地改善基于PCR的诊断工具有限的患者的初始筛查。
Since December 2019 the novel coronavirus SARS-CoV-2 has been identified as the cause of the pandemic COVID-19. Early symptoms overlap with other common conditions such as common cold and Influenza, making early screening and diagnosis are crucial goals for health practitioners. The aim of the study was to use machine learning (ML), an artificial neural network (ANN) and a simple statistical test to identify SARS-CoV-2 positive patients from full blood counts without knowledge of symptoms or history of the individuals. The dataset included in the analysis and training contains anonymized full blood counts results from patients seen at the Hospital Israelita Albert Einstein, at Sao Paulo, Brazil, and who had samples collected to perform the SARS-CoV-2 rt-PCR test during a visit to the hospital. Patient data was anonymised by the hospital, clinical data was standardized to have a mean of zero and a unit standard deviation. This data was made public with the aim to allow researchers to develop ways to enable the hospital to rapidly predict and potentially identify SARS-CoV-2 positive patients.We find that with full blood counts random forest, shallow learning and a flexible ANN model predict SARS-CoV-2 patients with high accuracy between populations on regular wards (AUC = 94-95%) and those not admitted to hospital or in the community (AUC = 80-86%). Here, AUC is the Area Under the receiver operating characteristics Curve and a measure for model performance. Moreover, a simple linear combination of 4 blood counts can be used to have an AUC of 85% for patients within the community. The normalised data of different blood parameters from SARS-CoV-2 positive patients exhibit a decrease in platelets, leukocytes, eosinophils, basophils and lymphocytes, and an increase in monocytes.SARS-CoV-2 positive patients exhibit a characteristic immune response profile pattern and changes in different parameters measured in the full blood count that are detected from simple and rapid blood tests. While symptoms at an early stage of infection are known to overlap with other common conditions, parameters of the full blood counts can be analysed to distinguish the viral type at an earlier stage than current rt-PCR tests for SARS-CoV-2 allow at present. This new methodology has potential to greatly improve initial screening for patients where PCR based diagnostic tools are limited.