Predicting binding affinities of emerging variants of SARS-CoV-2 using spike protein sequencing data: observations, caveats and recommendations

Predicting binding affinities of emerging variants of SARS-CoV-2 using spike protein sequencing data: observations, caveats and recommendations
复制标题

DOI:
10.1093/bib/bbac128
复制
发表时间:
2022-04-18
影响因子:
9.5
通讯作者:
Pal, Ranadip
Pal, Ranadip
中科院分区:
生物学2区
文献类型:
--
作者:
Zhang, Ruibo;Ghosh, Souparno;Pal, Ranadip

文献摘要

被引文献

相似文献

从氨基酸序列预测蛋白质性质是生物学和药理学中的一个重要问题。SARS-CoV-2刺突蛋白、人类受体和抗体之间的蛋白质-蛋白质相互作用是该病毒效力及其逃避人类免疫应答能力的关键决定因素。作为一种快速进化的病毒,SARS-CoV-2已经发展成许多变体,这些变体之间的毒力存在相当大的差异。因此,利用SARS-CoV-2的蛋白质组数据预测其病毒特性将大大有助于疾病的控制和预防。在本文中,我们回顾和比较了最近成功的预测方法,这些方法基于长短期记忆(LSTM),Transformer,卷积神经网络(CNN)和基于相似性的拓扑回归(TR)模型,并根据训练数据集和测试数据集之间的相似性提供了适当的预测方法的建议。我们比较了这些模型在预测SARS-CoV-2刺突蛋白序列的结合亲和力和表达方面的有效性。我们还探索了这些预测方法在实验室创建的数据上训练时的有效性,并负责预测从GISAID数据集获得的野生SARS-CoV-2刺突蛋白序列的结合亲和力。我们观察到,TR是一个更好的方法时,样本量小,测试蛋白质序列是足够相似的训练序列。然而,当训练样本量足够大并且预测需要外推时,LSTM嵌入和基于CNN的预测模型表现出上级性能。
Predicting protein properties from amino acid sequences is an important problem in biology and pharmacology. Protein-protein interactions among SARS-CoV-2 spike protein, human receptors and antibodies are key determinants of the potency of this virus and its ability to evade the human immune response. As a rapidly evolving virus, SARS-CoV-2 has already developed into many variants with considerable variation in virulence among these variants. Utilizing the proteomic data of SARS-CoV-2 to predict its viral characteristics will, therefore, greatly aid in disease control and prevention. In this paper, we review and compare recent successful prediction methods based on long short-term memory (LSTM), transformer, convolutional neural network (CNN) and a similarity-based topological regression (TR) model and offer recommendations about appropriate predictive methodology depending on the similarity between training and test datasets. We compare the effectiveness of these models in predicting the binding affinity and expression of SARS-CoV-2 spike protein sequences. We also explore how effective these predictive methods are when trained on laboratory-created data and are tasked with predicting the binding affinity of the in-the-wild SARS-CoV-2 spike protein sequences obtained from the GISAID datasets. We observe that TR is a better method when the sample size is small and test protein sequences are sufficiently similar to the training sequence. However, when the training sample size is sufficiently large and prediction requires extrapolation, LSTM embedding and CNN-based predictive model show superior performance.