In Silico Target Predictions: Defining a Benchmarking Data Set and Comparison of Performance of the Multiclass Naive Bayes and Parzen-Rosenblatt Window

In Silico Target Predictions: Defining a Benchmarking Data Set and Comparison of Performance of the Multiclass Naive Bayes and Parzen-Rosenblatt Window
复制标题

DOI:
10.1021/ci300435j
复制
发表时间:
2013-08-01
影响因子:
5.6
通讯作者:
Bender, Andreas
Bender, Andreas
中科院分区:
化学2区
文献类型:
--
作者:
Koutsoukas, Alexios;Lowe, Robert;Bender, Andreas

文献摘要

被引文献

相似文献

在这项研究中,比较了两种概率机器学习算法在生物活性分子的电子靶标预测中的应用,即著名的拉普拉斯改进的朴素贝叶斯分类器(NB)和最近引入的(化学信息学)Parzen-Rosenblatt窗口。这两个分类器都与从ChEMBL提取的生物活性化合物的大型数据集上的循环指纹一起进行了训练,这些数据集覆盖了894个人类蛋白质目标和超过155,000个配体-蛋白质对。由于其大小和包含的生物活性类别的数量,该数据集也被提供作为未来目标预测方法的基准数据集。除了对方法进行评估外,还探索了不同的绩效衡量标准。这不像在二进制分类设置中那样简单,这是由于类别的数量、多个类别成员的可能性以及将模型分数转换为评估模型性能的“是/否”预测的需要。在前1%的预测中,这两种算法的正确目标召回率都超过了80%。性能在很大程度上取决于给定类别的生物活性化合物的潜在多样性和大小,小类和低结构相似性对两种算法的影响程度不同。当在一个从袋熊中提取的外部测试集上进行测试时,通过排除所有与ChEMBL数据集中的化合物具有超过0.8的TAnimoto相似度的化合物,当前的方法在朴素贝叶斯和Parzen-Rosenblatt窗口的前1%中分别获得了63.3%和66.6%的召回率。虽然这些数字似乎表明性能较低,但对于需要为新的化学物质建立蛋白质目标的环境来说,它们也更现实。
In this study, two probabilistic machine-learning algorithms were compared for in silico target prediction of bioactive molecules, namely the well-established Laplacian-modified Naive Bayes classifier (NB) and the more recently introduced (to Cheminformatics) Parzen-Rosenblatt Window. Both classifiers were trained in conjunction with circular fingerprints on a large data set of bioactive compounds extracted from ChEMBL, covering 894 human protein targets with more than 155,000 ligand-protein pairs. This data set is also provided as a benchmark data set for future target prediction methods due to its size as well as the number of bioactivity classes it contains. In addition to evaluating the methods, different performance measures were explored. This is not as straightforward as in binary classification settings, due to the number of classes, the possibility of multiple class memberships, and the need to translate model scores into "yes/no" predictions for assessing model performance. Both algorithms achieved a recall of correct targets that exceeds 80% in the top 1% of predictions. Performance depends significantly on the underlying diversity and size of a given class of bioactive compounds, with small classes and low structural similarity affecting both algorithms to different degrees. When tested on an external test set extracted from WOMBAT covering more than 500 targets by excluding all compounds with Tanimoto similarity above 0.8 to compounds from the ChEMBL data set, the current methodologies achieved a recall of 63.3% and 66.6% among the top 1% for Naive Bayes and Parzen-Rosenblatt Window, respectively. While those numbers seem to indicate lower performance, they are also more realistic for settings where protein targets need to be established for novel chemical substances.