Pre-trained protein language model sheds new light on the prediction of Arabidopsis protein-protein interactions.

Pre-trained protein language model sheds new light on the prediction of Arabidopsis protein-protein interactions.
复制标题

DOI:
10.1186/s13007-023-01119-6
复制
发表时间:
2023-12-07
期刊:
影响因子:
5.1
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

蛋白质-蛋白质相互作用(PPIs)在许多生物过程中起着重要作用。因此,鉴定模式植物拟南芥中PPIs对深入了解植物生长发育,进而促进作物改良的基础研究具有重要意义。虽然目前已经确定了许多实验性拟南芥PPIs,但已知的拟南芥相互作用数据还远远不够完整。在此背景下,迫切需要从现有PPI数据中开发有效的机器学习模型来方便快速地预测未知的拟南芥PPI。我们使用一种名为ESM-1b的大规模预训练蛋白质语言模型(pLM)将蛋白质序列转换为高维向量,然后将其作为多层感知器(MLP)的输入。为了避免PPI预测中经常出现的性能高估,我们使用严格的数据集来训练和评估预测模型。结果表明,与其他plm或基线序列编码方案推断的预测模型相比,ESMAraPPI和ESM-1b的组合获得了更准确的性能。其中,在训练数据集中未见每对蛋白中的两个蛋白对的独立测试集上进行测试时,所提出的ESMAraPPI的AUPR值为0.810,表明其具有较强的泛化和外推能力。此外,提出的ESMAraPPI模型比几个最先进的通用或植物特异性PPI预测器表现更好。预训练模型ESM-1b的蛋白质序列嵌入包含丰富的蛋白质语义信息。通过与MLP算法的结合,ESM-1b在拟南芥ppi预测中表现出优异的性能。我们预计,所提出的预测模型(ESMAraPPI)可以作为一个非常有竞争力的工具,以加快鉴定拟南芥相互作用组。在线版本包含补充资料,网址为10.1186/s13007-023-01119-6。
Protein–protein interactions (PPIs) are heavily involved in many biological processes. Consequently, the identification of PPIs in the model plant Arabidopsis is of great significance to deeply understand plant growth and development, and then to promote the basic research of crop improvement. Although many experimental Arabidopsis PPIs have been determined currently, the known interactomic data of Arabidopsis is far from complete. In this context, developing effective machine learning models from existing PPI data to predict unknown Arabidopsis PPIs conveniently and rapidly is still urgently needed. We used a large-scale pre-trained protein language model (pLM) called ESM-1b to convert protein sequences into high-dimensional vectors and then used them as the input of multilayer perceptron (MLP). To avoid the performance overestimation frequently occurring in PPI prediction, we employed stringent datasets to train and evaluate the predictive model. The results showed that the combination of ESM-1b and MLP (i.e., ESMAraPPI) achieved more accurate performance than the predictive models inferred from other pLMs or baseline sequence encoding schemes. In particular, the proposed ESMAraPPI yielded an AUPR value of 0.810 when tested on an independent test set where both proteins in each protein pair are unseen in the training dataset, suggesting its strong generalization and extrapolating ability. Moreover, the proposed ESMAraPPI model performed better than several state-of-the-art generic or plant-specific PPI predictors. Protein sequence embeddings from the pre-trained model ESM-1b contain rich protein semantic information. By combining with the MLP algorithm, ESM-1b revealed excellent performance in predicting Arabidopsis PPIs. We anticipate that the proposed predictive model (ESMAraPPI) can serve as a very competitive tool to accelerate the identification of Arabidopsis interactome. The online version contains supplementary material available at 10.1186/s13007-023-01119-6.
DOI: 10.1038/nmeth.4083
发表时间: 2017-01
期刊: Nature methods
影响因子: 48
作者:
Li T;Wernersson R;Hansen RB;Horn H;Mercer J;Slodkowicz G;Workman CT;Rigina O;Rapacki K;Stærfeldt HH;Brunak S;Jensen TS;Lage K
通讯作者: Lage K
DOI: 10.1093/molbev/msx148
发表时间: 2017-08-01
影响因子: 10.7
作者:
Huerta-Cepas J;Forslund K;Coelho LP;Szklarczyk D;Jensen LJ;von Mering C;Bork P
通讯作者: Bork P
DOI: 10.1093/nar/gkw985
发表时间: 2017-01-04
影响因子: 14.9
作者:
Alanis-Lobato G;Andrade-Navarro MA;Schaefer MH
通讯作者: Schaefer MH
DOI: 10.1109/tpami.2021.3095381
发表时间: 2022-10-01
影响因子: 23.6
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1038/s41592-019-0598-1
发表时间: 2019-12-01
期刊: NATURE METHODS
影响因子: 48
作者:
Alley, Ethan C.;Khimulya, Grigory;Church, George M.
通讯作者: Church, George M.