Integrating experimental and literature protein-protein interaction data for protein complex prediction.

Integrating experimental and literature protein-protein interaction data for protein complex prediction.
复制标题

DOI:
10.1186/1471-2164-16-s2-s4
复制
发表时间:
2015
期刊:
影响因子:
4.4
通讯作者:
Wang J
Wang J
中科院分区:
生物学2区
文献类型:
--
作者:
Zhang Y;Lin H;Yang Z;Wang J

文献摘要

相似文献

蛋白质复合体的准确测定对于理解细胞的组织和功能至关重要。高通量的实验技术已经产生了大量的蛋白质-蛋白质相互作用(PPI)数据,使得从PPI网络预测蛋白质复合体成为可能。然而,高通量数据往往包括假阳性和假阴性,使得准确预测蛋白质复合体变得困难。生物医学文献包含大量的PPI数据,以及高通量的实验PPI数据,这些数据对于蛋白质复合体的预测是有价值的。在这项研究中,我们使用自然语言处理技术从生物医学文献中提取PPI数据。通过构建属性PPI网络,将这些数据与高通量的PPI和基因本体数据进行集成,提出了一种从属性PPI网络预测蛋白质复合体的新方法。该方法允许计算高通量和生物医学文献PPI数据的相对贡献。当应用于两个不同的酵母PPI数据集时,该方法能够准确地预测许多具有良好特征的蛋白质复合体。结果表明:(I)生物医学文献PPI数据能够有效提高蛋白质复合体预测的性能;(Ii)我们的方法充分利用了高通量的生物医学文献PPI数据和基因本体数据,实现了最先进的蛋白质复合体预测能力。
Accurate determination of protein complexes is crucial for understanding cellular organization and function. High-throughput experimental techniques have generated a large amount of protein-protein interaction (PPI) data, allowing prediction of protein complexes from PPI networks. However, the high-throughput data often includes false positives and false negatives, making accurate prediction of protein complexes difficult. The biomedical literature contains large quantities of PPI data that, along with high-throughput experimental PPI data, are valuable for protein complex prediction. In this study, we employ a natural language processing technique to extract PPI data from the biomedical literature. This data is subsequently integrated with high-throughput PPI and gene ontology data by constructing attributed PPI networks, and a novel method for predicting protein complexes from the attributed PPI networks is proposed. This method allows calculation of the relative contribution of high-throughput and biomedical literature PPI data. Many well-characterized protein complexes are accurately predicted by this method when apply to two different yeast PPI datasets. The results show that (i) biomedical literature PPI data can effectively improve the performance of protein complex prediction; (ii) our method makes good use of high-throughput and biomedical literature PPI data along with gene ontology data to achieve state-of-the-art protein complex prediction capabilities.