Improving compound-protein interaction prediction by building up highly credible negative samples.

Improving compound-protein interaction prediction by building up highly credible negative samples.
复制标题

DOI:
10.1093/bioinformatics/btv256
复制
发表时间:
2015-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Zhou S
Zhou S
中科院分区:
其他
文献类型:
--
作者:
Liu H;Sun J;Guan J;Zheng J;Zhou S

文献摘要

被引文献

相似文献

动机:化合物-蛋白质相互作用(CPI)的计算预测对于药物设计和开发非常重要,因为CPI的基因组规模实验验证不仅耗时而且昂贵。随着越来越多的有效相互作用的可用性,计算预测方法的性能受到缺乏可靠的负CPI样本的严重影响。一个系统的方法来筛选可靠的阴性样品成为关键,以提高性能的硅片预测方法。结果:本文旨在通过计算机筛选方法建立一套高可信度的CPI阴性样本。由于大多数现有的计算模型假设相似的化合物可能与相似的靶蛋白相互作用并实现显著的性能,因此基于与化合物的每个已知/预测靶标不相似的蛋白质不太可能被化合物靶向的匡威负命题来识别潜在的阴性样品是合理的,反之亦然。我们整合了各种资源,包括化合物的化学结构、化学表达谱和副作用、氨基酸序列、蛋白质-蛋白质相互作用网络和蛋白质的功能注释,形成了一个系统的筛选框架。我们首先在六个经典分类器上测试了筛选出的阴性样本,所有这些分类器在我们的阴性样本上的性能都明显高于随机生成的人类和秀丽隐杆线虫的阴性样本。然后我们在现有的三种预测模型上验证了负样本,包括二分局部模型,高斯核轮廓和贝叶斯矩阵分解,发现这些模型的性能在筛选的负样本上也有显著提高。此外,我们在药物生物活性数据集上验证了筛选的阴性样品。最后,我们通过在DrugBank中注释的积极相互作用和我们筛选的消极相互作用上训练支持向量机分类器来获得两组新的相互作用。筛选出的阴性样品和预测的相互作用为研究界提供了一个有用的资源,用于识别新的药物靶点,并对当前策划的化合物-蛋白质数据库进行了有益的补充。可用性:补充文件可在http://admis.fudan.edu.cn/negative-cpi/上获得。 联系方式:sgzhou@fudan.edu.cn补充信息:补充数据可在生物信息学在线获得。
Motivation: Computational prediction of compound–protein interactions (CPIs) is of great importance for drug design and development, as genome-scale experimental validation of CPIs is not only time-consuming but also prohibitively expensive. With the availability of an increasing number of validated interactions, the performance of computational prediction approaches is severely impended by the lack of reliable negative CPI samples. A systematic method of screening reliable negative sample becomes critical to improving the performance of in silico prediction methods. Results: This article aims at building up a set of highly credible negative samples of CPIs via an in silico screening method. As most existing computational models assume that similar compounds are likely to interact with similar target proteins and achieve remarkable performance, it is rational to identify potential negative samples based on the converse negative proposition that the proteins dissimilar to every known/predicted target of a compound are not much likely to be targeted by the compound and vice versa. We integrated various resources, including chemical structures, chemical expression profiles and side effects of compounds, amino acid sequences, protein–protein interaction network and functional annotations of proteins, into a systematic screening framework. We first tested the screened negative samples on six classical classifiers, and all these classifiers achieved remarkably higher performance on our negative samples than on randomly generated negative samples for both human and Caenorhabditis elegans. We then verified the negative samples on three existing prediction models, including bipartite local model, Gaussian kernel profile and Bayesian matrix factorization, and found that the performances of these models are also significantly improved on the screened negative samples. Moreover, we validated the screened negative samples on a drug bioactivity dataset. Finally, we derived two sets of new interactions by training an support vector machine classifier on the positive interactions annotated in DrugBank and our screened negative interactions. The screened negative samples and the predicted interactions provide the research community with a useful resource for identifying new drug targets and a helpful supplement to the current curated compound–protein databases. Availability: Supplementary files are available at: http://admis.fudan.edu.cn/negative-cpi/. Contact: sgzhou@fudan.edu.cn Supplementary Information: Supplementary data are available at Bioinformatics online.