Exploring the GDB-13 chemical space using deep generative models

Exploring the GDB-13 chemical space using deep generative models
复制标题

DOI:
10.1186/s13321-019-0341-z
复制
发表时间:
2019-03-12
影响因子:
8.6
通讯作者:
Engkvist, Ola
Engkvist, Ola
中科院分区:
化学2区
文献类型:
--
作者:
Arus-Pous, Josep;Blaschke, Thomas;Engkvist, Ola

文献摘要

被引文献

相似文献

最近的递归神经网络(RNN)的应用使得能够训练对化学空间进行采样的模型。在这项研究中,我们用分子串表示(SMILES)训练RNN,使用枚举数据库GDB-13(9.75亿个分子)的子集。我们发现,当对20亿个分子进行采样时,用100万个结构(数据库的0.1%)训练的模型在训练后再现了整个数据库的68.9%。我们还开发了一种使用负对数似然图来评估训练过程质量的方法。此外,我们使用了一个基于优惠券收集器问题的数学模型,该模型将训练后的模型与上限进行比较,因此我们能够量化它学到了多少。我们还建议,这种方法可以作为一种工具,基准的学习能力的任何分子生成模型架构。此外,生成的化学空间进行了分析,这表明,主要是由于SMILES的语法,复杂的分子与许多环和杂原子更难以采样。
Recent applications of recurrent neural networks (RNN) enable training models that sample the chemical space. In this study we train RNN with molecular string representations (SMILES) with a subset of the enumerated database GDB-13 (975 million molecules). We show that a model trained with 1 million structures (0.1% of the database) reproduces 68.9% of the entire database after training, when sampling 2 billion molecules. We also developed a method to assess the quality of the training process using negative log-likelihood plots. Furthermore, we use a mathematical model based on the coupon collector problem that compares the trained model to an upper bound and thus we are able to quantify how much it has learned. We also suggest that this method can be used as a tool to benchmark the learning capabilities of any molecular generative model architecture. Additionally, an analysis of the generated chemical space was performed, which shows that, mostly due to the syntax of SMILES, complex molecules with many rings and heteroatoms are more difficult to sample.