Cross-Modality Protein Embedding for Compound-Protein Affinity and Contact Prediction

Cross-Modality Protein Embedding for Compound-Protein Affinity and Contact Prediction
复制标题

DOI:
10.1101/2020.11.29.403162
复制
发表时间:
2020-11
期刊:
bioRxiv
影响因子:
--
通讯作者:
Yuning You;Yang Shen
Yuning You;Yang Shen
中科院分区:
其他
文献类型:
--
作者:
Yuning You;Yang Shen

文献摘要

相似文献

化合物-蛋白质对在 FDA 批准的药物-靶点对中占主导地位,化合物-蛋白质亲和力和接触 (CPAC) 的预测有助于加速药物发现。在这项研究中,我们将蛋白质视为多模式数据,包括一维氨基酸序列和(序列预测的)二维残基对接触图。我们凭经验评估了两种单一模式的嵌入在 CPAC 预测(即无结构可解释的化合物-蛋白质亲和力预测)的准确性和普遍性方面的表现。我们合理化了他们在嵌入个体模式和学习可概括的嵌入标签关系的挑战中的表现。我们进一步提出了两种涉及跨模态蛋白质嵌入的模型,并确定具有交叉相互作用的模型(从而捕获模态之间的相关性)在对训练集中从未见过的蛋白质的亲和力、接触和结合位点预测方面优于 SOTA 和我们的单一模态模型。
Compound-protein pairs dominate FDA-approved drug-target pairs and the prediction of compound-protein affinity and contact (CPAC) could help accelerate drug discovery. In this study we consider proteins as multi-modal data including 1D amino-acid sequences and (sequence-predicted) 2D residue-pair contact maps. We empirically evaluate the embeddings of the two single modalities in their accuracy and generalizability of CPAC prediction (i.e. structure-free interpretable compound-protein affinity prediction). And we rationalize their performances in both challenges of embedding individual modalities and learning generalizable embedding-label relationship. We further propose two models involving cross-modality protein embedding and establish that the one with cross interaction (thus capturing correlations among modalities) outperforms SOTAs and our single modality models in affinity, contact, and binding-site predictions for proteins never seen in the training set.