Classification and specific primer design for accurate detection of SARS-CoV-2 using deep learning.

Classification and specific primer design for accurate detection of SARS-CoV-2 using deep learning.
复制标题

基于深度学习的SARS-CoV-2分类及特异引物设计

DOI:
10.1038/s41598-020-80363-5
复制
发表时间:
2021-01-13
期刊:
影响因子:
4.6
通讯作者:
Kraneveld AD
Kraneveld AD
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lopez-Rincon A;Tonda A;Mendoza-Maldonado L;Mulders DGJC;Molenkamp R;Perez-Romero CA;Claassen E;Garssen J;Kraneveld AD

文献摘要

参考文献

被引文献

相似文献

在本文中,深度学习与可解释的人工智能技术相结合,发现了SARS-CoV-2的代表性基因组序列。首先对国家基因组学数据中心存储库中的553个序列进行卷积神经网络分类器训练,分离出冠状病毒科不同病毒株的基因组,准确率为98.73%。然后分析网络的行为,以发现模型用于识别SARS-CoV-2的序列,最终发现它独有的序列。发现的序列在来自国家生物技术信息中心和共享所有流感数据库全球倡议的样本上进行了验证,并被证明能够以近乎完美的准确性将SARS-CoV-2从不同的病毒株中分离出来。接下来,选择其中一个序列生成引物集,并与其他最先进的引物集进行测试,获得具有竞争力的结果。最后,合成引物并在患者样本(n = 6先前检测阳性)上进行测试,提供与常规诊断方法相似的灵敏度和100%的特异性。与现有方法相比,所提出的方法具有很大的附加价值,因为它既能够从有限的数据中自动识别有希望的病毒引物集,又能够在最短的时间内提供有效的结果。考虑到未来大流行的可能性,这些特征对于迅速制定具体的诊断检测方法非常宝贵。
In this paper, deep learning is coupled with explainable artificial intelligence techniques for the discovery of representative genomic sequences in SARS-CoV-2. A convolutional neural network classifier is first trained on 553 sequences from the National Genomics Data Center repository, separating the genome of different virus strains from the Coronavirus family with 98.73% accuracy. The network’s behavior is then analyzed, to discover sequences used by the model to identify SARS-CoV-2, ultimately uncovering sequences exclusive to it. The discovered sequences are validated on samples from the National Center for Biotechnology Information and Global Initiative on Sharing All Influenza Data repositories, and are proven to be able to separate SARS-CoV-2 from different virus strains with near-perfect accuracy. Next, one of the sequences is selected to generate a primer set, and tested against other state-of-the-art primer sets, obtaining competitive results. Finally, the primer is synthesized and tested on patient samples (n = 6 previously tested positive), delivering a sensitivity similar to routine diagnostic methods, and 100% specificity. The proposed methodology has a substantial added value over existing methods, as it is able to both automatically identify promising primer sets for a virus from a limited amount of data, and deliver effective results in a minimal amount of time. Considering the possibility of future pandemics, these characteristics are invaluable to promptly create specific detection methods for diagnostics.
DOI: 10.3390/v12111331
发表时间: 2020-11-19
期刊: Viruses
影响因子: --
作者:
Amoroso MG;Lucifora G;Degli Uberti B;Serra F;De Luca G;Borriello G;De Domenico A;Brandi S;Cuomo MC;Bove F;Riccardi MG;Galiero G;Fusco G
通讯作者: Fusco G
DOI: 10.1186/s12859-019-3050-8
发表时间: 2019-09-18
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Lopez-Rincon, Alejandro;Martinez-Archundia, Marlet;Tonda, Alberto
通讯作者: Tonda, Alberto
DOI: 10.1093/bib/bbt078
发表时间: 2014-05-01
影响因子: 9.5
作者:
Pinello, Luca;Lo Bosco, Giosue;Yuan, Guo-Cheng
通讯作者: Yuan, Guo-Cheng
DOI: 10.1099/vir.0.82856-0
发表时间: 2007-11-01
影响因子: 3.8
作者:
Padhan, Kartika;Tamar, Charu;Jameel, Shahid
通讯作者: Jameel, Shahid
DOI: 10.3390/cancers12071785
发表时间: 2020-07-01
期刊: CANCERS
影响因子: 5.2
作者:
Lopez-Rincon, Alejandro;Mendoza-Maldonado, Lucero;Tonda, Alberto
通讯作者: Tonda, Alberto