Virtual Screening for R-Groups, including Predicted pIC50 Contributions, within Large Structural Databases, Using Topomer CoMFA

Virtual Screening for R-Groups, including Predicted pIC50 Contributions, within Large Structural Databases, Using Topomer CoMFA
复制标题

DOI:
10.1021/ci8001556
复制
发表时间:
2008-11-01
影响因子:
5.6
通讯作者:
Soltanshahi, Farhad
Soltanshahi, Farhad
中科院分区:
化学2区
文献类型:
--
作者:
Cramer, Richard D.;Cruz, Phillip;Soltanshahi, Farhad

文献摘要

被引文献

相似文献

在填充大型结构数据库的大多数分子结构中,多个R基团(单价片段)是隐式可访问的。R基团搜索最好考虑pIC50贡献预测以及配体相似性或对接分数。然而,无论有没有pIC50预测,R组搜索目前都不实用。最普遍和最可靠的pIC50预测来源,即现有的3D-QSAR方法,也是困难和有点客观的。然而,在基于现场的3D-QSAR处理已经成功的25个数据集试验中,用客观(规范产生的)拓扑体姿势代替原始的结构引导的手动比对产生了可接受的3D-QSAR模型,平均而言具有与已发表模型几乎相同的统计质量,并且可以忽略不计的努力。他们的总pIC50预测误差为0.805,计算的是这25个拓扑体CoMFA模型的pIC50预测标准偏差的平均值,这些标准偏差来自1,109个可能的“遗漏一个R-基团”(LOORG)pIC50贡献。(这一新的LOORG协议提供了比通常的“遗漏一组化合物”LOO方法更现实和更严格的预测准确性测试。)相关的平均预测r(2)为0.495,表明PIC50预测精度大致介于完美和无用之间。为了评估基于Topmer-CoMFA的虚拟筛选识别“高活性”R基团的能力,采用了受试者工作曲线(ROC)方法。使用预测的pIC50值大于观察到的pIC50值范围的前25%作为“高活性”R基团的二元标准,25个拓扑体CoMFA模型的平均ROC面积为0.729。传统上解释的是,一个“高度活跃的”R-基团确实给予如此高的PIC50值的几率是0.729/(1-0.729)或几乎是3:1。为了确认在已实现的结构的大集合内的虚拟筛选将提供有用的数量和种类的R-基团建议,将形状相似性与“高活性”的PIC50相结合,将这25个模型提供的50个搜索应用于锌数据库内200,000个结构中的220,000个结构不同的R-基团候选,识别出每次搜索的平均5,705个R-基团,预测的最高pIC50组合平均比报道的最高pIC50大1.6个对数单位。
Multiple R-groups (monovalent fragments) are implicitly accessible within most of the molecular structures that populate large structural databases. R-group searching would desirably consider pIC50 contribution forecasts as well as ligand similarities or docking scores. However, R-group searching, with or without pIC50 forecasts, is currently not practical. The most prevalent and reliable source of pIC50 predictions, existing 3D-QSAR approaches, is also difficult and somewhat objective. Yet in 25 of 25 trials on data sets on which a field-based 3D-QSAR treatment had already succeeded, substitution of objective (canonically generated) topomer poses for the original structure-guided manual alignments produced acceptable 3D-QSAR models, on average having almost equivalent statistical quality to the published models, and with negligible effort. Their overall pIC50 prediction error is 0.805, calculated as the average over these 25 topomer CoMFA models in the standard deviations of pIC50 predictions, derived from the 1109 possible "leave-out-one-R-group" (LOORG) pIC50 contributions. (This novel LOORG protocol provides a more realistic and stringent test of prediction accuracy than the customary "leave-out-one-compound" Loo approach.) The associated average predictive r(2) of 0.495 indicates a pIC50 prediction accuracy roughly halfway between perfect and useless. To assess the ability of topomer-CoMFA based virtual screening to identify "highly active" R-groups, a Receiver Operating Curve (ROC) approach was adopted. Using, as the binary criterion for a "highly active" R-group, a predicted pIC50 greater than the top 25% of the observed pIC50 range, the ROC area averaged across the 25 topomer CoMFA models is 0.729. Conventionally interpreted, the odds that a "highly active" R-group will indeed confer such a high pIC50 are 0.729/(1-0.729) or almost 3 to 1. To confirm that virtual screening within large collections of realized structures would provide a useful quantity and variety of R-group suggestions, combining shape similarity with the "highly active" pIC50, the 50 searches provided by these 25 models were applied to 2.2 million structurally distinct R-group candidates among 2.0 million structures within a ZINC database, identifying an average of 5705 R-groups per search, with the highest predicted pIC50 combination averaging 1.6 log units greater than the highest reported pIC50s.