Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening

Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening
复制标题

DOI:
10.1371/journal.pone.0220113
复制
发表时间:
2019-08-20
期刊:
影响因子:
3.7
通讯作者:
Kurtzman, Tom
Kurtzman, Tom
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Chen, Lieyang;Cruz, Anthony;Kurtzman, Tom

文献摘要

被引文献

相似文献

最近,人们投入了大量精力,使用在蛋白质-配体复合物的 3D 结构图像上训练的卷积神经网络 (CNN) 模型来区分结合配体和非结合配体以进行虚拟筛选。然而,由于缺乏可靠的蛋白质-配体 X 射线结构和结合亲和力数据,需要使用构建的数据集来训练和评估 CNN 分子识别模型。在这里,我们概述了一种广泛使用的数据集“有用诱饵目录:增强型”(DUD-E) 中的各种偏差来源。我们构建并执行了测试,以调查使用 DUD-E 开发的 CNN 模型是否按照预期正确学习分子识别的基础物理,或者学习数据集本身固有的偏差。我们发现 CNN 模型中卓越的富集效率可归因于 DUD-E 数据集中隐藏的模拟和诱饵偏差,而不是蛋白质配体相互作用模式的成功概括。比较在 PDBbind 数据集上训练的其他深度学习模型,我们发现它们使用 DUD-E 的富集性能并不优于对接程序 AutoDock Vina 的性能。总之,这些结果表明,在将构建的数据集中可能存在的偏差应用于基于机器学习的方法开发之前,应该对其进行彻底评估。
Recently much effort has been invested in using convolutional neural network (CNN) models trained on 3D structural images of protein-ligand complexes to distinguish binding from non-binding ligands for virtual screening. However, the dearth of reliable protein-ligand x-ray structures and binding affinity data has required the use of constructed datasets for the training and evaluation of CNN molecular recognition models. Here, we outline various sources of bias in one such widely-used dataset, the Directory of Useful Decoys: Enhanced (DUD-E). We have constructed and performed tests to investigate whether CNN models developed using DUD-E are properly learning the underlying physics of molecular recognition, as intended, or are instead learning biases inherent in the dataset itself. We find that superior enrichment efficiency in CNN models can be attributed to the analogue and decoy bias hidden in the DUD-E dataset rather than successful generalization of the pattern of proteinligand interactions. Comparing additional deep learning models trained on PDBbind datasets, we found that their enrichment performances using DUD-E are not superior to the performance of the docking program AutoDock Vina. Together, these results suggest that biases that could be present in constructed datasets should be thoroughly evaluated before applying them to machine learning based methodology development.