Comparative Assessment of Scoring Functions on an Updated Benchmark: 1. Compilation of the Test Set

Comparative Assessment of Scoring Functions on an Updated Benchmark: 1. Compilation of the Test Set
复制标题

DOI:
10.1021/ci500080q
复制
发表时间:
2014-06-01
影响因子:
5.6
通讯作者:
Wang, Renxiao
Wang, Renxiao
中科院分区:
化学2区
文献类型:
--
作者:
Li, Yan;Liu, Zhihai;Wang, Renxiao

文献摘要

被引文献

相似文献

评分函数通常与分子对接方法结合应用,以预测配体结合位姿和配体结合亲和力或通过虚拟筛选鉴定活性化合物。一个客观的基准评估目前的评分功能的性能,预计将提供切实可行的指导,为用户作出明智的选择可用的方法。它还可以阐明当前方法的共同弱点,以便将来改进。我们的评分函数比较评估(CASF)项目的主要目标是提供一个高标准的,公开访问的这种类型的基准。我们最新的研究,即,CASF-2013评估了一组更新的蛋白质-配体复合物上的20种常用评分函数。该数据集是通过相当复杂的过程从PDBbind数据库(2013版)中记录的8302种蛋白质-配体复合物中选择的。通过考虑复杂结构的质量以及结合数据来进行样品选择。最后,合格的复合物在蛋白质序列中以90%的相似性进行聚类。从每个簇中选择三个代表性复合物以控制样品冗余。最终结果,即PDBbind核心集(2013版),由65个簇中的195个蛋白质配体复合物组成,结合常数跨越近10个数量级。在这个数据集中,82%的配体分子是“类药物”的,78%的蛋白质分子是经过验证的或潜在的药物靶点。讨论了配体的结合常数与几个关键性质之间的关系。评分函数评价的方法和结果将在本期的配套工作(doi:10.1021/ci 500081 m)中描述。
Scoring functions are often applied in combination with molecular docking methods to predict ligand binding poses and ligand binding affinities or to identify active compounds through virtual screening. An objective benchmark for assessing the performance of current scoring functions is expected to provide practical guidance for the users to make smart choices among available methods. It can also elucidate the common weakness in current methods for future improvements. The primary goal of our comparative assessment of scoring functions (CASF) project is to provide a high-standard, publicly accessible benchmark of this type. Our latest study, i.e., CASF-2013, evaluated 20 popular scoring functions on an updated set of protein-ligand complexes. This data set was selected out of 8302 protein-ligand complexes recorded in the PDBbind database (version 2013) through a fairly complicated process. Sample selection was made by considering the quality of complex structures as well as binding data. Finally, qualified complexes were clustered by 90% similarity in protein sequences. Three representative complexes were chosen from each cluster to control sample redundancy. The final outcome, namely, the PDBbind core set (version 2013), consists of 195 protein ligand complexes in 65 clusters with binding constants spanning nearly 10 orders of magnitude. In this data set, 82% of the ligand molecules are "druglike" and 78% of the protein molecules are validated or potential drug targets. Correlation between binding constants and several key properties of ligands are discussed. Methods and results of the scoring function evaluation will be described in a companion work in this issue (doi: 10.1021/ci500081m).