Community benchmarks for virtual screening

Community benchmarks for virtual screening
复制标题

DOI:
10.1007/s10822-008-9189-4
复制
发表时间:
2008-03-01
影响因子:
3.5
通讯作者:
Irwin, John J.
Irwin, John J.
中科院分区:
生物学3区
文献类型:
--
作者:
Irwin, John J.

文献摘要

被引文献

相似文献

在顶级命中中的配体富集是虚拟筛选的关键度量。为了避免偏见,诱饵应该类似的配体物理,使富集是不是由于简单的差异,总的功能。因此,我们创建了一个有用的诱饵(DUD)的目录,通过选择诱饵,类似的注释配体的物理,但不拓扑基准对接性能。DUD具有2950个注释配体和95,316个针对40个靶标的性质匹配的诱饵。据我所知,这是迄今为止规模最大、最全面的虚拟筛查项目基准公共数据集。本文概述了几种方法,DUD可以改进,以提供更好的遥测研究人员寻求了解目前的对接方法的优点和缺点。我还强调了粗心的几个陷阱:过度优化的风险,化学空间的问题,以及使用DUD的适当范围。仔细注意基准的组成以及如何使用它们对于避免被过度拟合和偏见误导至关重要。
Ligand enrichment among top-ranking hits is a key metric of virtual screening. To avoid bias, decoys should resemble ligands physically, so that enrichment is not attributable to simple differences of gross features. We therefore created a directory of useful decoys (DUD) by selecting decoys that resembled annotated ligands physically but not topologically to benchmark docking performance. DUD has 2950 annotated ligands and 95,316 property-matched decoys for 40 targets. It is by far the largest and most comprehensive public data set for benchmarking virtual screening programs that I am aware of. This paper outlines several ways that DUD can be improved to provide better telemetry to investigators seeking to understand both the strengths and the weaknesses of current docking methods. I also highlight several pitfalls for the unwary: a risk of over-optimization, questions about chemical space, and the proper scope for using DUD. Careful attention to both the composition of benchmarks and how they are used is essential to avoid being misled by overfitting and bias.