Improving (Q)SAR predictions by examining bias in the selection of compounds for experimental testing.

Improving (Q)SAR predictions by examining bias in the selection of compounds for experimental testing.
复制标题

通过检查实验测试化合物选择中的偏差来改进 (Q)SAR 预测。

DOI:
10.1080/1062936x.2019.1665580
复制
发表时间:
2019
影响因子:
3
通讯作者:
Poroikov,VV
Poroikov,VV
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Pogodin,PV;Lagunin,AA;Filimonov,DA;Nicklaus,MC;Poroikov,VV

文献摘要

相似文献

现有的关于结构和生物活性的数据有限,而且在不同的分子靶标和化合物中分布不均。如果这些数据代表了化学-生物相互作用总体的无偏见样本,问题就出现了。为了回答这个问题,我们分析了87,583个分子的ChEMBL数据,这些分子使用监督和非监督方法针对919个蛋白质靶标进行测试。使用化学开发工具包生成的Murcko框架的层次聚类显示,可用的数据形成了一个大的扩散云,没有明显的结构。与此形成对比的是,基于PASS的分类器允许预测化合物是否针对特定的分子靶标进行了测试,无论它是否具有活性。因此,人们可以得出结论,针对特定目标测试的化合物的选择是有偏见的,可能是由于先验知识的影响。我们利用这一事实评估了改善(Q)SAR预测的可能性:对于预测为对照目标进行测试的化合物,通过预测与特定目标的相互作用的准确性显著高于未测试的预测(平均ROC AUC分别约为0.87和0.75)。因此,考虑训练集数据中存在的偏差可能会提高虚拟筛选的性能。
Existing data on structures and biological activities are limited and distributed unevenly across distinct molecular targets and chemical compounds. The question arises if these data represent an unbiased sample of the general population of chemical-biological interactions. To answer this question, we analyzed ChEMBL data for 87,583 molecules tested against 919 protein targets using supervised and unsupervised approaches. Hierarchical clustering of the Murcko frameworks generated using Chemistry Development Toolkit showed that the available data form a big diffuse cloud without apparent structure. In contrast hereto, PASS-based classifiers allowed prediction whether the compound had been tested against the particular molecular target, despite whether it was active or not. Thus, one may conclude that the selection of chemical compounds for testing against specific targets is biased, probably due to the influence of prior knowledge. We assessed the possibility to improve (Q)SAR predictions using this fact: PASS prediction of the interaction with the particular target for compounds predicted as tested against the target has significantly higher accuracy than for those predicted as untested (average ROC AUC are about 0.87 and 0.75, respectively). Thus, considering the existing bias in the data of the training set may increase the performance of virtual screening.