From Alpha to Zeta: Identifying Variants and Subtypes of SARS-CoV-2 Via Clustering

From Alpha to Zeta: Identifying Variants and Subtypes of SARS-CoV-2 Via Clustering
复制标题

从 Alpha 到 Zeta:通过聚类识别 SARS-CoV-2 的变异体和亚型

DOI:
10.1089/cmb.2021.0302
复制
发表时间:
2021
影响因子:
1.7
通讯作者:
Patterson, Murray
Patterson, Murray
中科院分区:
生物学4区
文献类型:
--
作者:
Melnyk, Andrew;Mohebbi, Fatemeh;Knyazev, Sergey;Sahoo, Bikram;Hosseini, Roya;Skums, Pavel;Zelikovsky, Alex;Patterson, Murray

文献摘要

相似文献

GISAID(全球流感数据共享倡议)和 EMBL-EBI(欧洲分子生物学实验室 - 欧洲生物信息学研究所)(英国)等公共数据库中提供了数百万条 SARS-CoV-2(严重急性呼吸系统综合症-冠状病毒-2)序列,可以对病毒的进化、基因组多样性和动态进行前所未有的详细研究。在这里,我们通过采用最初为宿主内病毒群体单倍型分析而设计的方法对序列进行聚类来识别 SARS-CoV-2 的新变体和亚型。我们使用聚类熵来评估我们的结果——这是第一次在这种情况下使用它。与其他方法相比,我们的聚类方法达到了更低的熵,并且我们能够通过间隙填充和基于蒙特卡罗的熵最小化进一步提高这一点。此外,我们的方法清楚地识别了英国和 GISAID 数据集中众所周知的 Alpha 变体,并且还能够检测 GISAID 数据集中代表性较少(<1% 的序列)Beta(南非)、Epsilon(加利福尼亚)以及 Gamma 和 Zeta(巴西)变体。最后,我们表明,根据其簇随时间的增长速度,识别出的每个变体都具有高选择性适应性。这表明我们的聚类方法是检测非常大的数据集中罕见的子类型的可行替代方法。
The availability of millions of SARS-CoV-2 (Severe Acute Respiratory Syndrome-Coronavirus-2) sequences in public databases such as GISAID (Global Initiative on Sharing All Influenza Data) and EMBL-EBI (European Molecular Biology Laboratory-European Bioinformatics Institute) (the United Kingdom) allows a detailed study of the evolution, genomic diversity, and dynamics of a virus such as never before. Here, we identify novel variants and subtypes of SARS-CoV-2 by clustering sequences in adapting methods originally designed for haplotyping intrahost viral populations. We asses our results using clustering entropy—the first time it has been used in this context. Our clustering approach reaches lower entropies compared with other methods, and we are able to boost this even further through gap filling and Monte Carlo-based entropy minimization. Moreover, our method clearly identifies the well-known Alpha variant in the U.K. and GISAID data sets, and is also able to detect the much less represented (<1% of the sequences) Beta (South Africa), Epsilon (California), and Gamma and Zeta (Brazil) variants in the GISAID data set. Finally, we show that each variant identified has high selective fitness, based on the growth rate of its cluster over time. This demonstrates that our clustering approach is a viable alternative for detecting even rare subtypes in very large data sets.