Cross-modality and self-supervised protein embedding for compound–protein affinity and contact prediction

Cross-modality and self-supervised protein embedding for compound–protein affinity and contact prediction
复制标题

用于化合物-蛋白质亲和力和接触预测的跨模态和自监督蛋白质嵌入

DOI:
10.1093/bioinformatics/btac470
复制
发表时间:
2022
期刊:
影响因子:
5.8
通讯作者:
Shen, Yang
Shen, Yang
中科院分区:
生物学3区
文献类型:
--
作者:
You, Yuning;Shen, Yang

文献摘要

相似文献

化合物-蛋白质亲和和接触(CPAC)预测的计算方法旨在通过同时预测化合物-蛋白质相互作用的强度和模式来促进合理的药物发现。虽然期望的输出高度依赖于结构,但缺乏蛋白质结构往往使无结构方法仅依赖于蛋白质序列输入。具有亲和力和接触标签的化合物-蛋白质对的稀缺性进一步限制了CPAC模型的准确性和通用性。为了克服上述结构天真和标记数据稀缺的挑战,我们分别引入了跨模态和自监督学习,用于结构感知和任务相关的蛋白质嵌入。具体来说,蛋白质数据可以以一维氨基酸序列和预测的二维接触图的两种方式获得,这些接触图分别分别嵌入循环神经网络和图形神经网络,以及联合嵌入两种交叉模态方案。此外,通过利用大量未标记的蛋白质数据,两种蛋白质模式都在各种自监督学习策略下进行了预训练。我们的研究结果表明,个体蛋白质模式在预测亲和或接触的强度上有所不同。适当的跨模态蛋白质嵌入结合自监督学习,在预测未知蛋白质的亲和力和接触时提高了模型的泛化性。可用性和实施数据和源代码可在https://github.com/Shen-Lab/CPAC.Supplementary上获得。补充数据可在bioinformatics online上获得。
MotivationComputational methods for compound–protein affinity and contact (CPAC) prediction aim at facilitating rational drug discovery by simultaneous prediction of the strength and the pattern of compound–protein interactions. Although the desired outputs are highly structure-dependent, the lack of protein structures often makes structure-free methods rely on protein sequence inputs alone. The scarcity of compound–protein pairs with affinity and contact labels further limits the accuracy and the generalizability of CPAC models.ResultsTo overcome the aforementioned challenges of structure naivety and labeled-data scarcity, we introduce cross-modality and self-supervised learning, respectively, for structure-aware and task-relevant protein embedding. Specifically, protein data are available in both modalities of 1D amino-acid sequences and predicted 2D contact maps that are separately embedded with recurrent and graph neural networks, respectively, as well as jointly embedded with two cross-modality schemes. Furthermore, both protein modalities are pre-trained under various self-supervised learning strategies, by leveraging massive amount of unlabeled protein data. Our results indicate that individual protein modalities differ in their strengths of predicting affinities or contacts. Proper cross-modality protein embedding combined with self-supervised learning improves model generalizability when predicting both affinities and contacts for unseen proteins.Availability and implementationData and source codes are available at https://github.com/Shen-Lab/CPAC.Supplementary informationSupplementary data are available atBioinformaticsonline.