Pitfalls of machine learning models for protein-protein interactions

Pitfalls of machine learning models for protein-protein interactions
复制标题

DOI:
10.1101/2022.02.07.479382
复制
发表时间:
2022-02
期刊:
bioRxiv
影响因子:
--
通讯作者:
Loïc Lannelongue;M. Inouye
Loïc Lannelongue;M. Inouye
中科院分区:
其他
文献类型:
--
作者:
Loïc Lannelongue;M. Inouye

文献摘要

被引文献

相似文献

蛋白质-蛋白质相互作用(PPI)对于理解生物途径及其在发育和疾病中的作用至关重要。基于经典机器学习的计算工具在预测Silico中的PPI方面取得了成功,但这项任务缺乏一致和可靠的框架,导致网络模型难以比较,算法之间的差异仍然无法解释。为了更好地理解支撑这些模型的潜在推理机制,我们设计了一个开源的基准测试框架,该框架在促进重复性的同时,考虑了一系列生物学和统计陷阱。我们用它来阐明网络拓扑的影响,以及不同的算法如何处理高度连接的蛋白质。通过研究基于功能基因组学和基于序列的人类PPI模型,我们证明了它们的互补性,因为前者在单个蛋白质上表现最好,而后者专门研究涉及中枢的相互作用。我们还表明,算法设计对功能基因组数据的性能影响很小。我们在人类和酿酒酵母数据之间复制了我们的结果,并证明使用功能基因组学的模型更适合跨物种的PPI预测。随着序列和功能基因组学数据的迅速增加,我们的研究为未来PPI网络的构建、比较和应用提供了原则性的基础。
Protein-protein interactions (PPIs) are essential to understanding biological pathways as well as their roles in development and disease. Computational tools, based on classic machine learning, have been successful at predicting PPIs in silico, but the lack of consistent and reliable frameworks for this task has led to network models that are difficult to compare and discrepancies between algorithms that remain unexplained. To better understand the underlying inference mechanisms that underpin these models, we designed an open-source framework for benchmarking that accounts for a range of biological and statistical pitfalls while facilitating reproducibility. We use it to shed light on the impact of network topology and how different algorithms deal with highly connected proteins. By studying functional genomics-based and sequence-based models on human PPIs, we show their complementarity as the former performs best on lone proteins while the latter specialises in interactions involving hubs. We also show that algorithm design has little impact on performance with functional genomic data. We replicate our results between both human and S. cerevisiae data and demonstrate that models using functional genomics are better suited to PPI prediction across species. With rapidly increasing amounts of sequence and functional genomics data, our study provides a principled foundation for future construction, comparison and application of PPI networks.