A Weak Supervised Learning Method for Essential Protein Detection Based on STRING Database and Learning Representation

A Weak Supervised Learning Method for Essential Protein Detection Based on STRING Database and Learning Representation
复制标题

一种基于STRING数据库和学习表示的弱监督学习必需蛋白检测方法

DOI:
10.1109/bibm.2018.8621469
复制
发表时间:
2018
期刊:
2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
影响因子:
--
通讯作者:
Hongfei Lin
Hongfei Lin
中科院分区:
--
文献类型:
--
作者:
Zhizheng Wang;Yuanyuan Sun;Yawen Guan;Yibin Zhang;Liang Yang;Kan Xu;Yijia Zhang;Hongfei Lin

文献摘要

被引文献

相似文献

蛋白质-蛋白质相互作用(PPI)网络中必需蛋白质的检测对于了解生物体的功能非常重要。目前,用于必需蛋白质搜索的算法主要基于网络拓扑结构和先验知识,因此忽略了PPI网络本身所包含的生物学知识。为此,我们提出了整合相互作用蛋白质数据库和STRING数据库的关键蛋白质搜索算法(IDSSP),优先搜索的蛋白质由STRING数据库中得分最高的蛋白质组成。此外,我们提出了一个弱监督学习算法的基础上IDSSP的结果。我们首先对IDSSP算法结果中的必需蛋白质进行标记。然后,我们利用表示学习算法和STRING数据库提取PPI网络节点的特征。最后,使用机器学习分类算法对必需蛋白质进行分类。对蛋白质的搜索结果表明,IDSSP算法在最佳情况下的top-k精度比现有方法提高了11.9%。对必需蛋白质分类的结果表明,结合生物学特征的分类方法的F1得分高于仅结合拓扑学特征的分类方法。因此,充分利用STRING数据库所包含的生物学信息,比单纯利用拓扑特征更有效地进行必需蛋白质的检测。
The detection of essential proteins in the protein-protein interaction (PPI) network is important for understanding the functions of organisms. At present, the algorithms used for essential protein search are mainly based on network topology and prior knowledge, so the biological knowledge contained in the PPI network itself is neglected. Therefore, we proposed the algorithm that integrates the Database of Interacting Proteins and STRING database to search essential proteins (IDSSP), in which prior proteins are composed of the highest-scoring proteins in STRING database. In addition, we propose a weak supervised learning algorithm based on the results of IDSSP. We label the essential proteins in the IDSSP algorithm results at first. Then, we extract features of the PPI network nodes by utilizing the representation learning algorithm and STRING database. Finally, the machine learning classification algorithms are used to classify the essential proteins. The results of searching essential proteins show that the top-k precision of IDSSP algorithm has increased by 11.9% compared with the state-of-art methods in the best situation. The results on essential protein classification indicate that the F1-score of classification methods combining with biological features are higher than those with only topological features. In a conclusion, making full use of biological information contained by STRING database is more effective than only using topological features in the task of essential protein detection.