Computationally Probing Drug-Protein Interactions Via Support Vector Machine

Computationally Probing Drug-Protein Interactions Via Support Vector Machine
复制标题

通过支持向量机计算探测药物-蛋白质相互作用

DOI:
10.2174/157018010791163433
复制
发表时间:
2010-06-01
影响因子:
1
通讯作者:
Deng, Nai-Yang
Deng, Nai-Yang
中科院分区:
医学4区
文献类型:
--
作者:
Wang, Yong-Cui;Yang, Zhi-Xia;Deng, Nai-Yang

文献摘要

被引文献

相似文献

在过去的几十年里,由于小分子(药物、代谢物或配体)和蛋白质之间的物理和遗传相互作用的规模和复杂性,人们对它们之间的关系进行了广泛的研究。特别是,计算预测药物与蛋白质的相互作用对于加快新型治疗药物的开发进程至关重要。本文通过引入两种机器学习思想,提出了一种用于药物-蛋白质相互作用预测的有监督学习方法--支持向量机。首先,将药物间的化学结构相似性和蛋白质间的基因组序列相似性直观地编码为特征向量来表示给定的药物-蛋白质对。其次,针对训练数据不均衡的问题,设计了一种自动选取金标正值数据集的过程,即金标正值数据相对于大规模未标注数据是稀缺的。我们的基于支持向量机的预测模型在四类药物靶蛋白上得到了验证,包括酶、离子通道、G蛋白偶联受体和核受体。我们发现,我们的方法改进了现有方法在给定的假阳性率的情况下的真阳性率。功能标注分析和数据库搜索表明,我们的新预测值得进一步的实验验证。此外,后续分析表明,我们的方法可以部分捕捉药物-蛋白质相互作用网络的拓扑特征。综上所述,我们的新方法可以有效地识别潜在的药物-蛋白质结合,并将促进药物发现的进一步研究。
The past decades witnessed extensive efforts to study the relationships among small molecules (drugs, metabolites, or ligands) and proteins due to the scale and complexity of their physical and genetic interactions. Particularly, computationally predicting the drug-protein interactions is fundamentally important in speeding up the process of developing novel therapeutic agents. Here, we present a supervised learning method, support vector machine (SVM), to predict drug-protein interactions by introducing two machine learning ideas. Firstly, the chemical structure similarity among drugs and the genomic sequence similarity among proteins are intuitively encoded as a feature vector to represent a given drug-protein pair. Secondly, we design an automatic procedure to select a gold-standard negative dataset to deal with the training data imbalance issue, i.e., gold-standard positive data is scarce relative to large scale unlabeled data. Our SVM based predictor is validated on four classes of drug target proteins, including enzymes, ion channels, G-protein couple receptors, and nuclear receptors. We find that our method improves the existing methods regarding to true positive rate upon given false positive rate. The functional annotation analysis and database search indicate that our new predictions are worthy of future experimental validation. In addition, follow-up analysis suggests that our method can partly capture the topological features in the drug-protein interaction network. In conclusion, our new method can efficiently identify the potential drug-protein bindings and will promote the further research in drug discovery.