How to Achieve Better Results Using PASS-Based Virtual Screening: Case Study for Kinase Inhibitors.

How to Achieve Better Results Using PASS-Based Virtual Screening: Case Study for Kinase Inhibitors.
复制标题

DOI:
10.3389/fchem.2018.00133
复制
发表时间:
2018
影响因子:
5.5
通讯作者:
Poroikov VV
Poroikov VV
中科院分区:
化学3区
文献类型:
--
作者:
Pogodin PV;Lagunin AA;Rudik AV;Filimonov DA;Druzhilovskiy DS;Nicklaus MC;Poroikov VV

文献摘要

相似文献

目前,利用可合成的虚拟库存(SAVI)库的可能性推动了新药物的发现,该库包括约2.83亿个分子,每个分子都标有从商业上可获得的起始材料的拟议合成一步路线。SAVI数据库非常适合基于配体的虚拟筛选方法,以选择用于实验测试的分子。在这项研究中,我们比较了三种分析结构-活性关系的方法的性能,这些方法在选择包括在训练集中的“活性”和“非活性”化合物的标准上有所不同。PASS(物质活性光谱预测)是基于改进的朴素贝叶斯算法的,因为它已被证明是健壮的,即使训练集中的信息不完整,也能仅基于化合物的结构式提供对许多生物活性的良好预测。我们在这个案例研究中使用了不同的激酶抑制剂亚组,因为目前关于这类重要的类药物分子的许多数据都是可用的。基于从ChEMBL20数据库中提取的激酶抑制剂的子集,我们进行了PASS训练,然后将该模型应用于ChEMBL20中尚未存在的ChEMBL23化合物,以识别新的激酶抑制剂。正如人们可能预期的那样,如果只使用在训练过程中针对不同的激酶的实验确认的活性和非活性化合物,则获得最好的预测精度。然而,对于一些激酶,即使我们使用合并的训练集,我们也获得了合理的结果,在这些训练集中,我们将没有针对特定激酶进行测试的化合物指定为非活性化合物。因此,根据特定生物活性的数据的可用性,人们可以选择第一种或第二种方法来创建基于配体的计算工具,以在虚拟筛选中获得可能的最佳结果。
Discovery of new pharmaceutical substances is currently boosted by the possibility of utilization of the Synthetically Accessible Virtual Inventory (SAVI) library, which includes about 283 million molecules, each annotated with a proposed synthetic one-step route from commercially available starting materials. The SAVI database is well-suited for ligand-based methods of virtual screening to select molecules for experimental testing. In this study, we compare the performance of three approaches for the analysis of structure-activity relationships that differ in their criteria for selecting of “active” and “inactive” compounds included in the training sets. PASS (Prediction of Activity Spectra for Substances), which is based on a modified Naïve Bayes algorithm, was applied since it had been shown to be robust and to provide good predictions of many biological activities based on just the structural formula of a compound even if the information in the training set is incomplete. We used different subsets of kinase inhibitors for this case study because many data are currently available on this important class of drug-like molecules. Based on the subsets of kinase inhibitors extracted from the ChEMBL 20 database we performed the PASS training, and then applied the model to ChEMBL 23 compounds not yet present in ChEMBL 20 to identify novel kinase inhibitors. As one may expect, the best prediction accuracy was obtained if only the experimentally confirmed active and inactive compounds for distinct kinases in the training procedure were used. However, for some kinases, reasonable results were obtained even if we used merged training sets, in which we designated as inactives the compounds not tested against the particular kinase. Thus, depending on the availability of data for a particular biological activity, one may choose the first or the second approach for creating ligand-based computational tools to achieve the best possible results in virtual screening.