CSAR benchmark exercise of 2010: selection of the protein-ligand complexes.

CSAR benchmark exercise of 2010: selection of the protein-ligand complexes.
复制标题

DOI:
10.1021/ci200082t
复制
发表时间:
2011-09-26
影响因子:
5.6
通讯作者:
Carlson HA
Carlson HA
中科院分区:
化学2区
文献类型:
--
作者:
Dunbar JB Jr;Smith RD;Yang CY;Ung PM;Lexa KW;Khazanov NA;Stuckey JA;Wang S;Carlson HA

文献摘要

被引文献

相似文献

药物设计的一个主要目标是改进对接和评分的计算方法。社区结构活动资源(CSAR)旨在从行业和学术界收集可用于此目的的可用数据()。此外,CSAR还负责根据收集的数据组织全社区的活动。这些练习中的第一个旨在使用蛋白质-配体复合物的大型且多样化的数据集来衡量对接和评分的总体状态。要求参与者计算所提供的复合物的亲和力,然后重新计算可能改进其特定方法的变化。该第一数据集选自现有PDB条目,其在Binding MOAD中具有结合数据(Kd或Ki),用来自PDB bind的条目扩增。最终的数据集包含343种不同的蛋白质-配体复合物,跨度为14 pKd。16种蛋白质在数据集中具有三个或更多个复合物,用户可以从该复合物开始检查同源系列。固有的实验误差限制了分数和测量的亲和力之间可能的相关性;当拟合到数据集而没有过度参数化时,R2被限制为0.9。当使用在外部数据上训练的方法对数据集进行评分时,R2被限制为10.8。详细介绍了数据集最初是如何选择的,以及它成熟的过程,以更好地满足社区的需求。许多团体慷慨地参与了改进数据集的工作,这突出了支持性的、协作性的努力在推动我们的领域向前发展方面的价值。
A major goal in drug design is the improvement of computational methods for docking and scoring. The Community Structure Activity Resource (CSAR) aims to collect available data from industry and academia which may be used for this purpose (). Also, CSAR is charged with organizing community-wide exercises based on the collected data. The first of these exercises was aimed to gauge the overall state of docking and scoring, using a large and diverse data set of protein–ligand complexes. Participants were asked to calculate the affinity of the complexes as provided and then recalculate with changes which may improve their specific method. This first data set was selected from existing PDB entries which had binding data (Kd or Ki) in Binding MOAD, augmented with entries from PDBbind. The final data set contains 343 diverse protein–ligand complexes and spans 14 pKd. Sixteen proteins have three or more complexes in the data set, from which a user could start an inspection of congeneric series. Inherent experimental error limits the possible correlation between scores and measured affinity; R2 is limited to ∼0.9 when fitting to the data set without over parametrizing. R2 is limited to ∼0.8 when scoring the data set with a method trained on outside data. The details of how the data set was initially selected, and the process by which it matured to better fit the needs of the community are presented. Many groups generously participated in improving the data set, and this underscores the value of a supportive, collaborative effort in moving our field forward.