Comparative Assessment of Scoring Functions on an Updated Benchmark: 2. Evaluation Methods and General Results

Comparative Assessment of Scoring Functions on an Updated Benchmark: 2. Evaluation Methods and General Results
复制标题

DOI:
10.1021/ci500081m
复制
发表时间:
2014-06-01
影响因子:
5.6
通讯作者:
Wang, Renxiao
Wang, Renxiao
中科院分区:
化学2区
文献类型:
--
作者:
Li, Yan;Han, Li;Wang, Renxiao

文献摘要

被引文献

相似文献

我们的评分函数比较评估(CASF)基准是为了提供对当前评分函数的客观评估而创建的。CASF的关键思想是比较评分功能在不同蛋白质配体复合物上的一般性能。为了避免在分子对接的背景下测试评分函数,通过使用先前生成的配体结合姿态集合将评分过程从对接(或采样)过程中分离出来。在这里,我们描述了最新的CASF-2013研究的技术方法和评估结果。本研究采用PDBbind核心集(2013版)作为主要测试集,该核心集由195个具有高质量三维结构和可靠结合常数的蛋白质-配体复合物组成。我们从“评分能力”(结合亲和力预测)、“排名能力”(相对排名预测)、“对接能力”(结合姿态预测)和“筛选能力”(从随机分子中区分真正的结合物)四个方面对20个评分功能进行了评估,这些评分功能大多在主流商业软件中实现。我们的研究结果表明,这些评分函数在对接/筛选能力测试中的表现通常比在评分/排名能力测试中的表现更有希望。评分能力测试中排名靠前的评分功能,如X-Score(HM)、ChemScore@SYBYL、ChemPLP@GOLD、PLP@DS等,也在评分能力测试中排名靠前。在对接功率测试中排名靠前的评分功能,如ChemPLP@GOLD、Chemscore@GOLD、GlidScore-SP、LigScore@DS、PLP@DS等,也在筛选功率测试中排名靠前。我们在整个测试集及其子集上获得的结果表明,蛋白质-配体结合亲和力预测的真正挑战在于极性相互作用和相关的脱溶效应。在高亲和力的蛋白质配体复合物中观察到的非加性特征也需要注意。
Our comparative assessment of scoring functions (CASF) benchmark is created to provide an objective evaluation of current scoring functions. The key idea of CASF is to compare the general performance of scoring functions on a diverse set of protein-ligand complexes. In order to avoid testing scoring functions in the context of molecular docking, the scoring process is separated from the docking (or sampling) process by using ensembles of ligand binding poses that are generated in prior. Here, we describe the technical methods and evaluation results of the latest CASF-2013 study. The PDBbind core set (version 2013) was employed as the primary test set in this study, which consists of 195 protein-ligand complexes with high-quality three-dimensional structures and reliable binding constants. A panel of 20 scoring functions, most of which are implemented in main-stream commercial software, were evaluated in terms of "scoring power" (binding affinity prediction), "ranking power" (relative ranking prediction), "docking power" (binding pose prediction), and "screening power" (discrimination of true binders from random molecules). Our results reveal that the performance of these scoring functions is generally more promising in the docking/screening power tests than in the scoring/ranking power tests. Top-ranked scoring functions in the scoring power test, such as X-Score(HM), ChemScore@SYBYL, ChemPLP@GOLD, and PLP@DS, are also top-ranked in the ranking power test. Top-ranked scoring functions in the docking power test, such as ChemPLP@GOLD, Chemscore@GOLD, GlidScore-SP, LigScore@DS, and PLP@DS, are also top-ranked in the screening power test. Our results obtained on the entire test set and its subsets suggest that the real challenge in protein-ligand binding affinity prediction lies in polar interactions and associated desolvation effect. Nonadditive features observed among high-affinity protein-ligand complexes also need attention.