Machine Learning on DNA-Encoded Libraries: A New Paradigm for Hit Finding

Machine Learning on DNA-Encoded Libraries: A New Paradigm for Hit Finding
复制标题

DOI:
10.1021/acs.jmedchem.0c00452
复制
发表时间:
2020-08-27
影响因子:
7.3
通讯作者:
Riley, Patrick
Riley, Patrick
中科院分区:
医学1区
文献类型:
--
作者:
McCloskey, Kevin;Sigel, Eric A.;Riley, Patrick

文献摘要

被引文献

相似文献

DNA编码的小分子文库(DEL)已经能够发现许多具有治疗价值的不同蛋白质靶点的新型抑制剂。我们展示了一种将机器学习应用于DEL选择数据的新方法,通过从大型商业和易于合成的化合物库中识别活性分子。我们仅使用DEL选择数据训练模型,并将自动或可自动化的过滤器应用于预测。我们对三种不同的蛋白质靶点进行了一项大型前瞻性研究(类似于2000种化合物):sEH(水解酶),ER α(核受体)和c-KIT(激酶)。该方法是有效的,在30 μ M时的总体命中率接近30%,并且发现了针对每个靶标的有效化合物(IC 50 < 10 nM)。即使对于与原始DEL不同的分子,该系统也能做出有用的预测,并且识别的化合物多种多样,主要是药物样的,并且与已知的配体不同。这项工作展示了一个强大的新方法来命中发现。
DNA-encoded small molecule libraries (DELs) have enabled discovery of novel inhibitors for many distinct protein targets of therapeutic value. We demonstrate a new approach applying machine learning to DEL selection data by identifying active molecules from large libraries of commercial and easily synthesizable compounds. We train models using only DEL selection data and apply automated or automatable filters to the predictions. We perform a large prospective study (similar to 2000 compounds) across three diverse protein targets: sEH (a hydrolase), ER alpha (a nuclear receptor), and c-KIT (a kinase). The approach is effective, with an overall hit rate of similar to 30% at 30 mu M and discovery of potent compounds (IC50 < 10 nM) for every target. The system makes useful predictions even for molecules dissimilar to the original DEL, and the compounds identified are diverse, predominantly drug-like, and different from known ligands. This work demonstrates a powerful new approach to hit-finding.