Classifying "kinase inhibitor-likeness" by using machine-learning methods

Classifying "kinase inhibitor-likeness" by using machine-learning methods
复制标题

DOI:
10.1002/cbic.200400109
复制
发表时间:
2005-03-01
期刊:
影响因子:
3.2
通讯作者:
Günther, J
Günther, J
中科院分区:
生物学3区
文献类型:
--
作者:
Briem, H;Günther, J

文献摘要

被引文献

相似文献

通过使用由 Ghose-Crippen 参数编码的内部小分子结构数据集,应用多种机器学习技术来区分激酶抑制剂和其他没有报告对任何蛋白激酶活性的分子。所有四种方法——支持向量机 (SVM)、人工神经网络 (ANN)、具有 GA 优化特征选择的 k 最近邻分类 (GAANN) 和递归分区 (RP)——都被证明能够提供合理的区分。然而,观察到这些方法之间的性能存在显着差异。对于所有测试的技术,使用派生的 13 种不同模型的共识投票提高了预测的准确性、精确度、召回率和 F1 值方面的质量。在比较各个模型的平均值时,支持向量机以及 GA/kNN 组合的表现优于其他技术。通过使用各自的多数票,神经网络的预测产生了最高的 F1 值,其次是支持向量机。
By using an in-house data set of small-molecule structures, encoded by Ghose-Crippen parameters, several machine learning techniques were applied to distinguish between kinase inhibitors and other molecules with no reported activity on any protein kinase. All four approaches pursued-support-vector machines (SVM), artificial neural networks (ANN), k nearest neighbor classification with GA-optimized feature selection (GAANN), and recursive partitioning (RP)-proved capable of providing a reasonable discrimination. Nevertheless, substantial differences in performance among the methods were observed. For all techniques tested, the use of a consensus vote of the 13 different models derived improved the quality of the predictions in terms of accuracy, precision, recall, and F1 value. Support-vector machines, followed by the GA/kNN combination, outperformed the other techniques when comparing the average of individual models. By using the respective majority votes, the prediction of neural networks yielded the highest F1 value, followed by SVMs.