Genome-wide analysis of SARS-CoV-2 virus strains circulating worldwide implicates heterogeneity

Genome-wide analysis of SARS-CoV-2 virus strains circulating worldwide implicates heterogeneity
复制标题

DOI:
10.1038/s41598-020-70812-6
复制
发表时间:
2020-08-19
期刊:
影响因子:
4.6
通讯作者:
Hossain, M. Anwar
Hossain, M. Anwar
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Islam, M. Rafiul;Hoque, M. Nazmul;Hossain, M. Anwar

文献摘要

被引文献

相似文献

严重急性呼吸综合征冠状病毒-2 (SARS-CoV-2)是一种新型进化发散型RNA病毒,是当前破坏性的COVID-19大流行的罪魁祸首。为了探索基因组特征,我们全面分析了截至2020年3月30日全球报告给GISAID数据库的2492个完整和/或接近完整的SARS-CoV-2菌株基因组序列。全基因组注释显示,在整个SARS-CoV-2基因组的不同位置存在1,516个核苷酸水平的变异。此外,核苷酸(nt)缺失分析发现,除了先前报道的ORF8(开放阅读框)、spike和ORF7a蛋白编码序列缺失外,整个基因组中还有12个缺失位点,特别是在多蛋白ORF1ab (n=9)、ORF10 (n=1)和3 '-UTR (n=2)中。来自系统基因水平突变和蛋白质谱分析的证据显示大量氨基酸(aa)替换(n=744),表明病毒蛋白异质。值得注意的是,除了上海分离株hCoV-19/Shanghai/SH0007/2020 (EPI_ISL_416320)的隐表位378位赖氨酸被精氨酸取代外,与血管紧张素转换酶2 (ACE2)和交叉反应中和抗体相互作用的受体结合域(RBD)残基在分析的病毒株中都是保守的。此外,我们对SARS-CoV-2感染的初步流行病学数据结果显示,欧洲(43.07%)、亚洲(38.09%)和北美(29.64%)的SARS-CoV-2基因组序列的aa突变频率相对较高,而欧洲温带国家(如意大利、西班牙、荷兰、法国、英国和比利时)的病死率仍然较高。因此,目前在大流行早期阶段采用的基因组注释方法可能是一种很有前途的工具,用于监测和跟踪不断演变的大流行情况、相关的遗传变异及其对制定有效控制和预防战略的影响。
Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2), a novel evolutionary divergent RNA virus, is responsible for the present devastating COVID-19 pandemic. To explore the genomic signatures, we comprehensively analyzed 2,492 complete and/or near-complete genome sequences of SARS-CoV-2 strains reported from across the globe to the GISAID database up to 30 March 2020. Genome-wide annotations revealed 1,516 nucleotide-level variations at different positions throughout the entire genome of SARS-CoV-2. Moreover, nucleotide (nt) deletion analysis found twelve deletion sites throughout the genome other than previously reported deletions at coding sequence of the ORF8 (open reading frame), spike, and ORF7a proteins, specifically in polyprotein ORF1ab (n=9), ORF10 (n=1), and 3 '-UTR (n=2). Evidence from the systematic gene-level mutational and protein profile analyses revealed a large number of amino acid (aa) substitutions (n=744), demonstrating the viral proteins heterogeneous. Notably, residues of receptor-binding domain (RBD) showing crucial interactions with angiotensin-converting enzyme 2 (ACE2) and cross-reacting neutralizing antibody were found to be conserved among the analyzed virus strains, except for replacement of lysine with arginine at 378th position of the cryptic epitope of a Shanghai isolate, hCoV-19/Shanghai/SH0007/2020 (EPI_ISL_416320). Furthermore, our results of the preliminary epidemiological data on SARS-CoV-2 infections revealed that frequency of aa mutations were relatively higher in the SARS-CoV-2 genome sequences of Europe (43.07%) followed by Asia (38.09%), and North America (29.64%) while case fatality rates remained higher in the European temperate countries, such as Italy, Spain, Netherlands, France, England and Belgium. Thus, the present method of genome annotation employed at this early pandemic stage could be a promising tool for monitoring and tracking the continuously evolving pandemic situation, the associated genetic variants, and their implications for the development of effective control and prophylaxis strategies.