Improved K-mer Based Prediction of Protein-Protein Interactions With Chaos Game Representation, Deep Learning and Reduced Representation Bias

Improved K-mer Based Prediction of Protein-Protein Interactions With Chaos Game Representation, Deep Learning and Reduced Representation Bias
复制标题

DOI:
10.48550/arxiv.2310.14764
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Ruth Veevers;Dan MacLean
Ruth Veevers;Dan MacLean
中科院分区:
其他
文献类型:
--
作者:
Ruth Veevers;Dan MacLean

文献摘要

相似文献

蛋白质-蛋白质相互作用驱动许多生物过程,包括通过植物 R 蛋白和细胞表面受体检测植物病原体。许多机器学习研究试图预测蛋白质-蛋白质相互作用,但性能高度依赖于训练数据;当所涉及的蛋白质包含在训练数据中时,模型已被证明可以准确预测相互作用,但当应用于以前未见过的蛋白质时,其结果始终较差。此外,使用参与多种相互作用的蛋白质训练的模型可能会受到表示偏差的影响,其中预测不是由学习到的生物特征驱动的,而是由学习相互作用数据集的结构驱动的。我们提出了一种从交互数据集中提取唯一对的方法,为无偏机器学习生成非冗余配对数据。将该方法应用于包含拟南芥和病原体效应子相互作用的数据集后,我们开发了一种卷积神经网络模型,能够从蛋白质编码基因的混沌游戏表示中学习和预测相互作用。
Protein-protein interactions drive many biological processes, including the detection of phytopathogens by plants' R-Proteins and cell surface receptors. Many machine learning studies have attempted to predict protein-protein interactions but performance is highly dependent on training data; models have been shown to accurately predict interactions when the proteins involved are included in the training data, but achieve consistently poorer results when applied to previously unseen proteins. In addition, models that are trained using proteins that take part in multiple interactions can suffer from representation bias, where predictions are driven not by learned biological features but by learning of the structure of the interaction dataset. We present a method for extracting unique pairs from an interaction dataset, generating non-redundant paired data for unbiased machine learning. After applying the method to datasets containing _Arabidopsis thaliana_ and pathogen effector interations, we developed a convolutional neural network model capable of learning and predicting interactions from Chaos Game Representations of proteins' coding genes.