Similarity-Based Methods and Machine Learning Approaches for Target Prediction in Early Drug Discovery: Performance and Scope

Similarity-Based Methods and Machine Learning Approaches for Target Prediction in Early Drug Discovery: Performance and Scope
复制标题

基于相似性的方法和机器学习方法用于早期药物发现中的靶点预测:性能和范围

DOI:
10.3390/ijms21103585
复制
发表时间:
2020-05-01
影响因子:
5.6
通讯作者:
Kirchmair, Johannes
Kirchmair, Johannes
中科院分区:
生物学2区
文献类型:
--
作者:
Mathai, Neann;Kirchmair, Johannes

文献摘要

被引文献

相似文献

用于预测药物和类药物化合物的大分子靶标的计算方法已经发展成为药物发现的关键技术。然而,已建立的验证协议留下了一些关键问题的性能和方法的范围未得到解决。例如,预测成功率通常被报告为测试集的所有化合物的平均值,并且不考虑单个测试化合物与训练实例之间的结构关系。为了更好地理解基于配体的方法对目标预测的价值,我们在三种测试场景下对基于相似性的方法和基于随机森林的机器学习方法(均采用2D分子指纹)进行了基准测试:使用外部数据的标准测试场景,标准时间分割场景,以及设计为最接近真实世界条件的场景。此外,我们根据单个测试分子与训练数据的距离对结果进行了去卷积。我们发现,令人惊讶的是,基于相似性的方法在所有测试场景中的表现通常优于机器学习方法,即使在查询在结构上与训练(或参考)数据中的实例明显不同的情况下,尽管已知目标空间的覆盖率要高得多。
Computational methods for predicting the macromolecular targets of drugs and drug-like compounds have evolved as a key technology in drug discovery. However, the established validation protocols leave several key questions regarding the performance and scope of methods unaddressed. For example, prediction success rates are commonly reported as averages over all compounds of a test set and do not consider the structural relationship between the individual test compounds and the training instances. In order to obtain a better understanding of the value of ligand-based methods for target prediction, we benchmarked a similarity-based method and a random forest based machine learning approach (both employing 2D molecular fingerprints) under three testing scenarios: a standard testing scenario with external data, a standard time-split scenario, and a scenario that is designed to most closely resemble real-world conditions. In addition, we deconvoluted the results based on the distances of the individual test molecules from the training data. We found that, surprisingly, the similarity-based approach generally outperformed the machine learning approach in all testing scenarios, even in cases where queries were structurally clearly distinct from the instances in the training (or reference) data, and despite a much higher coverage of the known target space.