Large-scale comparison of machine learning methods for drug target prediction on ChEMBL.

Large-scale comparison of machine learning methods for drug target prediction on ChEMBL.
复制标题

DOI:
10.1039/c8sc00148k
复制
发表时间:
2018-06-28
期刊:
影响因子:
8.4
通讯作者:
Hochreiter S
Hochreiter S
中科院分区:
化学1区
文献类型:
--
作者:
Mayr A;Klambauer G;Unterthiner T;Steijaert M;Wegner JK;Ceulemans H;Clevert DA;Hochreiter S

文献摘要

参考文献

被引文献

相似文献

迄今为止,对9种最先进的药物靶点预测方法进行的最大规模的比较研究发现,深度学习优于所有其他竞争对手。这些结果是基于1300种检测和50万种化合物的基准。深度学习是目前广泛应用领域中最成功的机器学习技术,最近已成功应用于药物发现研究,以预测潜在的药物靶点并筛选活性分子。然而,由于(1)缺乏大规模研究,(2)药物发现数据集特有的化合物系列偏差,以及(3)大量潜在深度学习架构带来的超参数选择偏差,目前尚不清楚深度学习是否真的可以在药物发现任务中超越现有的计算方法。因此,我们评估了几种深度学习方法在大规模药物发现数据集上的性能,并将结果与其他机器学习和目标预测方法的结果进行了比较。为了避免超参数选择或复合系列的潜在偏差,我们使用了嵌套聚类交叉验证策略。我们发现(1)深度学习方法明显优于所有竞争方法,(2)深度学习的预测性能在许多情况下与湿实验室中进行的测试相当(即,体外测定)。
The to date largest comparative study of nine state-of-the-art drug target prediction methods finds that deep learning outperforms all other competitors. The results are based on a benchmark of 1300 assays and half a million compounds. Deep learning is currently the most successful machine learning technique in a wide range of application areas and has recently been applied successfully in drug discovery research to predict potential drug targets and to screen for active molecules. However, due to (1) the lack of large-scale studies, (2) the compound series bias that is characteristic of drug discovery datasets and (3) the hyperparameter selection bias that comes with the high number of potential deep learning architectures, it remains unclear whether deep learning can indeed outperform existing computational methods in drug discovery tasks. We therefore assessed the performance of several deep learning methods on a large-scale drug discovery dataset and compared the results with those of other machine learning and target prediction methods. To avoid potential biases from hyperparameter selection or compound series, we used a nested cluster-cross-validation strategy. We found (1) that deep learning methods significantly outperform all competing methods and (2) that the predictive performance of deep learning is in many cases comparable to that of tests performed in wet labs (i.e., in vitro assays).
DOI: 10.1021/ci010132r
发表时间: 2002-11-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
Durant, JL;Leland, BA;Nourse, JG
通讯作者: Nourse, JG
ChemoPy:用于计算生物学和化学信息学的免费 Python 包
DOI: 10.1093/bioinformatics/btt105
发表时间: 2013-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cao, Dong-Sheng;Xu, Qing-Song;Liang, Yi-Zeng
通讯作者: Liang, Yi-Zeng
DOI: 10.1023/a:1007379606734
发表时间: 1997-07-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Caruana, R
通讯作者: Caruana, R
DOI: 10.1093/nar/gkt1031
发表时间: 2014-01
影响因子: 14.9
作者:
Bento AP;Gaulton A;Hersey A;Bellis LJ;Chambers J;Davies M;Krüger FA;Light Y;Mak L;McGlinchey S;Nowotka M;Papadatos G;Santos R;Overington JP
通讯作者: Overington JP
DOI: 10.1186/s13321-014-0047-1
发表时间: 2014
影响因子: 8.6
作者:
Baumann D;Baumann K
通讯作者: Baumann K