Comparison of Random Forest and Pipeline Pilot Naive Bayes in Prospective QSAR Predictions

Comparison of Random Forest and Pipeline Pilot Naive Bayes in Prospective QSAR Predictions
复制标题

DOI:
10.1021/ci200615h
复制
发表时间:
2012-03-01
影响因子:
5.6
通讯作者:
Voigt, Johannes H.
Voigt, Johannes H.
中科院分区:
化学2区
文献类型:
--
作者:
Chen, Bin;Sheridan, Robert P.;Voigt, Johannes H.

文献摘要

被引文献

相似文献

随机森林被认为是目前最好的QSAR方法之一,在预测的准确性。然而,它是计算密集型的。Naive Bayes是一种简单、鲁棒的分类方法。Laplacian修改的Naive Bayes实现是广泛使用的商业化学信息学平台Pipeline Pilot中的首选QSAR方法。我们比较了Pipeline Pilot Naive Bayes(PLPNB)和随机森林对18个大型、多样化的内部QSAR数据集进行准确预测的能力。这些活动包括达标活动和ADME相关活动。这些数据集被设置为二进制或多类别活动的分类问题。我们使用了一种划分训练集和测试集的时间分割方法,因为我们认为这是模拟前瞻性预测的一种现实方式。PLPNB在计算上是高效的。然而,随机森林预测至少和PLPNB在我们的数据集上的预测一样好,而且在许多情况下明显好于PLPNB。PLPNB使用ECFP 4和ECFP 6描述符(这些描述符是Pipeline Pilot原生的)时性能更好,而使用我们尝试的其他描述符时性能更差。
Random forest is currently considered one of the best QSAR methods available in terms of accuracy of prediction. However, it is computationally intensive. Naive Bayes is a simple, robust classification method. The Laplacian-modified Naive Bayes implementation is the preferred QSAR method in the widely used commercial chemoinformatics platform Pipeline Pilot. We made a comparison of the ability of Pipeline Pilot Naive Bayes (PLPNB) and random forest to make accurate predictions on 18 large, diverse in-house QSAR data sets. These include on-target and ADME-related activities. These data sets were set up as classification problems with either binary or multicategory activities. We used a time-split method of dividing training and test sets, as we feel this is a realistic way of simulating prospective prediction. PLPNB is computationally efficient. However, random forest predictions are at least as good and in many cases significantly better than those of PLPNB on our data sets. PLPNB performs better with ECFP4 and ECFP6 descriptors, which are native to Pipeline Pilot, and more poorly with other descriptors we tried.