Does Inter-Protein Contact Prediction Benefit from Multi-Modal Data and Auxiliary Tasks?

Does Inter-Protein Contact Prediction Benefit from Multi-Modal Data and Auxiliary Tasks?
复制标题

DOI:
10.1101/2022.11.29.518454
复制
发表时间:
2022-12
期刊:
bioRxiv
影响因子:
--
通讯作者:
Arghamitra Talukder;Rujie Yin;Yuanfei Sun;Yang Shen;Yuning You
Arghamitra Talukder;Rujie Yin;Yuanfei Sun;Yang Shen;Yuning You
中科院分区:
其他
文献类型:
--
作者:
Arghamitra Talukder;Rujie Yin;Yuanfei Sun;Yang Shen;Yuning You

文献摘要

相似文献

AlphaFold 2已经彻底改变了蛋白质结构的计算机预测方法,而那些预测蛋白质之间界面的方法相对不发达,因为蛋白质-蛋白质复合物的数据过于复杂但相对有限。简而言之,蛋白质是折叠成3D结构的1D氨基酸序列,并相互作用以形成功能组件。我们认为,这种复杂的情况下,更好地模拟与额外的指示性信息,反映其多模态性质和多尺度功能。为了改善蛋白质间残基-残基接触的二进制预测,我们建议用多模态表示来增强输入特征,并将目标与辅助预测任务协同。(i)我们首先逐步添加三种蛋白质模式到模型中:蛋白质序列,序列与进化信息,和结构感知的蛋白质内残基接触图。我们观察到,利用所有的数据模式提供了最好的预测精度。分析表明,进化和结构的信息有利于预测的困难和刚性的蛋白质复合物,分别评估的相似性,本地残基接触结合复杂的结构。(ii)接下来,我们通过自我监督的预训练(蛋白质-蛋白质相互作用(PPI)的二进制预测)和多任务学习(蛋白质间残基距离和角度的预测)引入三个辅助任务。虽然据报道,PPI预测受益于预测相互接触(作为因果解释),但在我们的研究中没有发现反之亦然。同样,更细粒度的距离和角度预测似乎也没有均匀地改善接触预测。这再次反映了蛋白质-蛋白质复合物数据的高度复杂性,因此设计和整合协同辅助任务仍然具有挑战性。
Approaches to in silico prediction of protein structures have been revolutionized by AlphaFold2, while those to predict interfaces between proteins are relatively underdeveloped, owing to the overly complicated yet relatively limited data of protein–protein complexes. In short, proteins are 1D sequences of amino acids folding into 3D structures, and interact to form assemblies to function. We believe that such intricate scenarios are better modeled with additional indicative information that reflects their multi-modality nature and multi-scale functionality. To improve binary prediction of inter-protein residue-residue contacts, we propose to augment input features with multi-modal representations and to synergize the objective with auxiliary predictive tasks. (i) We first progressively add three protein modalities into models: protein sequences, sequences with evolutionary information, and structure-aware intra-protein residue contact maps. We observe that utilizing all data modalities delivers the best prediction precision. Analysis reveals that evolutionary and structural information benefit predictions on the difficult and rigid protein complexes, respectively, assessed by the resemblance to native residue contacts in bound complex structures. (ii) We next introduce three auxiliary tasks via self-supervised pre-training (binary prediction of protein-protein interaction (PPI)) and multi-task learning (prediction of inter-protein residue–residue distances and angles). Although PPI prediction is reported to benefit from predicting inter-contacts (as causal interpretations), it is not found vice versa in our study. Similarly, the finer-grained distance and angle predictions did not appear to uniformly improve contact prediction either. This again reflects the high complexity of protein–protein complex data, for which designing and incorporating synergistic auxiliary tasks remains challenging.