A Weak Supervised Learning Method for Essential Protein Detection Based on STRING Database and Learning Representation
A Weak Supervised Learning Method for Essential Protein Detection Based on STRING Database and Learning Representation
复制标题
一种基于STRING数据库和学习表示的弱监督学习必需蛋白检测方法
DOI:
10.1109/bibm.2018.8621469
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Hongfei Lin
中科院分区:
文献类型:
--
作者:
Zhizheng Wang;Yuanyuan Sun;Yawen Guan;Yibin Zhang;Liang Yang;Kan Xu;Yijia Zhang;Hongfei Lin
The detection of essential proteins in the protein-protein interaction (PPI) network is important for understanding the functions of organisms. At present, the algorithms used for essential protein search are mainly based on network topology and prior knowledge, so the biological knowledge contained in the PPI network itself is neglected. Therefore, we proposed the algorithm that integrates the Database of Interacting Proteins and STRING database to search essential proteins (IDSSP), in which prior proteins are composed of the highest-scoring proteins in STRING database. In addition, we propose a weak supervised learning algorithm based on the results of IDSSP. We label the essential proteins in the IDSSP algorithm results at first. Then, we extract features of the PPI network nodes by utilizing the representation learning algorithm and STRING database. Finally, the machine learning classification algorithms are used to classify the essential proteins. The results of searching essential proteins show that the top-k precision of IDSSP algorithm has increased by 11.9% compared with the state-of-art methods in the best situation. The results on essential protein classification indicate that the F1-score of classification methods combining with biological features are higher than those with only topological features. In a conclusion, making full use of biological information contained by STRING database is more effective than only using topological features in the task of essential protein detection.