DrugE-Rank: improving drug-target interaction prediction of new candidate drugs or targets by ensemble learning to rank.

DrugE-Rank: improving drug-target interaction prediction of new candidate drugs or targets by ensemble learning to rank.
复制标题

DrugE-Rank:通过集成学习排序来改进新候选药物或靶标的药物-靶点相互作用预测

DOI:
10.1093/bioinformatics/btw244
复制
发表时间:
2016-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Zhu S
Zhu S
中科院分区:
其他
文献类型:
--
作者:
Yuan Q;Gao J;Wu D;Zhang S;Mamitsuka H;Zhu S

文献摘要

被引文献

相似文献

动机:确定药物-靶标相互作用是药物发现中的一项重要任务。为了减少大量的实验时间和财务成本,已经提出了许多计算方法。尽管这些方法使用了许多不同的原理,但它们的性能远不令人满意,特别是在预测新候选药物或靶点的药物-靶点相互作用方面。方法:基于机器学习的方法可以分为两类:基于特征的方法和基于相似性的方法。学习排序是基于特征的方法中最强大的技术。基于相似性的方法被广泛接受,因为它们的想法是连接化学和基因组空间,分别由药物和靶标相似性表示。我们提出了一种新的方法,DrugE-Rank,通过很好地结合两种不同类型的方法的优点来提高预测性能。也就是说,DrugE-Rank使用LTR,其中多个众所周知的基于相似性的方法可以用作集成学习的组件。结果如下:DrugE-Rank的性能通过使用DrugBank数据的三个主要实验进行了全面检查:(i)2014年3月之前FDA(美国食品和药物管理局)批准的药物的交叉验证;(ii)2014年3月之后FDA批准的药物的独立测试;以及(iii)FDA实验药物的独立测试。实验结果表明,DrugE-Rank算法的性能明显优于其他同类算法,尤其是对FDA批准的新药和FDA实验性药物,其预测召回率曲线下的面积(Area under Prediction Recall curve)提高了30%以上。可用性:http://datamining-iip.fudan.edu.cn/service/DrugE-Rank联系:zhusf@fudan.edu.cn补充信息:补充数据可在生物信息学在线。
Motivation: Identifying drug–target interactions is an important task in drug discovery. To reduce heavy time and financial cost in experimental way, many computational approaches have been proposed. Although these approaches have used many different principles, their performance is far from satisfactory, especially in predicting drug–target interactions of new candidate drugs or targets. Methods: Approaches based on machine learning for this problem can be divided into two types: feature-based and similarity-based methods. Learning to rank is the most powerful technique in the feature-based methods. Similarity-based methods are well accepted, due to their idea of connecting the chemical and genomic spaces, represented by drug and target similarities, respectively. We propose a new method, DrugE-Rank, to improve the prediction performance by nicely combining the advantages of the two different types of methods. That is, DrugE-Rank uses LTR, for which multiple well-known similarity-based methods can be used as components of ensemble learning. Results: The performance of DrugE-Rank is thoroughly examined by three main experiments using data from DrugBank: (i) cross-validation on FDA (US Food and Drug Administration) approved drugs before March 2014; (ii) independent test on FDA approved drugs after March 2014; and (iii) independent test on FDA experimental drugs. Experimental results show that DrugE-Rank outperforms competing methods significantly, especially achieving more than 30% improvement in Area under Prediction Recall curve for FDA approved new drugs and FDA experimental drugs. Availability: http://datamining-iip.fudan.edu.cn/service/DrugE-Rank Contact: zhusf@fudan.edu.cn Supplementary information: Supplementary data are available at Bioinformatics online.