Pros and cons of virtual screening based on public "Big Data": In silico mining for new bromodomain inhibitors

Pros and cons of virtual screening based on public "Big Data": In silico mining for new bromodomain inhibitors
复制标题

DOI:
10.1016/j.ejmech.2019.01.010
复制
发表时间:
2019-03-01
影响因子:
6.7
通讯作者:
Varnek, Alexandre
Varnek, Alexandre
中科院分区:
医学1区
文献类型:
--
作者:
Casciuc, Iuri;Horvath, Dragos;Varnek, Alexandre

文献摘要

被引文献

相似文献

本文描述的虚拟筛选(VS)研究旨在检测新的溴结构域BRD4结合,并依赖于来自公共数据库(ChEMBL,Reaxys)的知识来建立BRD活性的一组预测模型,用于推测配体的电子选择。除了实际发现新的BRD配体外,这还提供了一个机会,可以实际评估公共领域“大数据”对稳健预测模型建立的实际有用性。获得的模型被用来虚拟筛选来自Enamine公司集合的200万个化合物的集合。然后,这个工业伙伴对VS程序选择的2992个分子的子集进行了实验筛选,以确定它们具有很高的活性可能性。经过实验测试后,检测到29个确认命中,占选定候选者的1%。总而言之,这项研究再次强调,公共结构-活性数据库是当今药物发现的关键资产。然而,到目前为止,通过已发表的研究获得的最先进的知识限制了它们的实用性。靶点特异性构效信息很少足够丰富,其异质性使其在合理药物设计中的开发变得极其困难。此外,已发表的用于建立选择要进行实验筛选的化合物的模型的亲和力测量可能与实验命中选择标准没有很好的相关性(在实践中,通常是由设备限制强加的)。然而,与同等的随机筛选活动相比,机器学习的命中率强劲地提高了2.6倍,这表明机器学习能够提取一些真实的知识,尽管结构活性数据中存在所有噪音。(C)2019爱思唯尔·马森SAS。版权所有。
The Virtual Screening (VS) study described herein aimed at detecting novel Bromodomain BRD4 binders and relied on knowledge from public databases (ChEMBL, REAXYS) to establish a battery of predictive models of BRD activity for in silico selection of putative ligands. Beyond the actual discovery of new BRD ligands, this represented an opportunity to practically estimate the actual usefulness of public domain "Big Data" for robust predictive model building. Obtained models were used to virtually screen a collection of 2 million compounds from the Enamine company collection. This industrial partner then experimentally screened a subset of 2992 molecules selected by the VS procedure for their high likelihood to be active. Twenty nine confirmed hits were detected after experimental testing, representing 1% of the selected candidates. As a general conclusion, this study emphasizes once more that public structure-activity databases are nowadays key assets in drug discovery. Their usefulness is however limited by the state-of-the-art knowledge harvested so far by published studies. Target-specific structure-activity information is rarely rich enough, and its heterogeneity makes it extremely difficult to exploit in rational drug design. Furthermore, published affinity measures serving to build models selecting compounds to be experimentally screened may not be well correlated with the experimental hit selection criterion (in practice, often imposed by equipment constraints). Nevertheless, a robust 2.6 fold increase in hit rate with respect to an equivalent, random screening campaign showed that machine learning is able to extract some real knowledge in spite of all the noise in structure-activity data. (C) 2019 Elsevier Masson SAS. All rights reserved.