Open Source Bayesian Models. 2. Mining a "Big Dataset" To Create and Validate Models with ChEMBL

Open Source Bayesian Models. 2. Mining a "Big Dataset" To Create and Validate Models with ChEMBL
复制标题

DOI:
10.1021/acs.jcim.5b00144
复制
发表时间:
2015-06-01
影响因子:
5.6
通讯作者:
Ekins, Sean
Ekins, Sean
中科院分区:
化学2区
文献类型:
--
作者:
Clark, Alex M.;Ekins, Sean

文献摘要

被引文献

相似文献

在相关论文中,我们描述了使用扩展连接 (ECFP) 和最大直径 6 (FCFP) 型指纹的分子功能类指纹构建拉普拉斯校正朴素贝叶斯模型的参考实现。作为后续行动,我们现在已经进行了大规模的验证研究,以确保该技术能够推广到各种药物发现数据集。为了实现这一目标,我们使用了 ChEMBL(版本 20)数据库,并将其分成 2000 多个单独的数据集,每个数据集都包含具有相同目标和活性测量值的化合物和测量值。为了使用两种状态贝叶斯分类来测试这些数据集,我们开发了一种自动算法来检测活动/非活动指定的合适阈值,并将其应用于所有集合。通过这些数据集,我们能够确定我们的贝叶斯模型实现对于大多数情况都是有效的,并且我们能够量化指纹折叠对接收者算子曲线交叉验证指标的影响。我们还能够研究训练/测试集划分的选择对最终召回率的影响。这些数据集以及相应的模型数据文件已公开可供下载,可与 CDK 和多个移动应用程序结合使用。我们还探索了一些新颖的可视化方法,这些方法利用 ECFP/FCFP 指纹的结构起源来归属对活性有积极和消极贡献的分子区域。在跨生物体的数千个相关数据集中对分子进行评分的能力也可能有助于获得期望和不期望的脱靶效应,并为源自表型筛选的化合物提出潜在靶标。
In an associated paper, we have described a reference implementation of Laplacian-corrected naive Bayesian model building using extended connectivity (ECFP)- and molecular function class fingerprints of maximum diameter 6 (FCFP)-type fingerprints. As a follow-up, we have now undertaken a large-scale validation study in order to ensure that the technique generalizes to a broad variety of drug discovery datasets. To achieve this, we have used the ChEMBL (version 20) database and split it into more than 2000 separate datasets, each of which consists of compounds and measurements with the same target and activity measurement. In order to test these datasets with the two-state Bayesian classification, we developed an automated algorithm for detecting a suitable threshold for active/inactive designation, which we applied to all collections. With these datasets, we were able to establish that our Bayesian model implementation is effective for the large majority of cases, and we were able to quantify the impact of fingerprint folding on the receiver operator curve cross-validation metrics. We were also able to study the impact that the choice of training/testing set partitioning has on the resulting recall rates. The datasets have been made publicly available to be downloaded, along with the corresponding model data files, which can be used in conjunction with the CDK and several mobile apps. We have also explored some novel visualization methods which leverage the structural origins of the ECFP/FCFP fingerprints to attribute regions of a molecule responsible for positive and negative contributions to activity. The ability to score molecules across thousands of relevant datasets across organisms also may help to access desirable and undesirable off-target effects as well as suggest potential targets for compounds derived from phenotypic screens.