Drug Analogs from Fragment-Based Long Short-Term Memory Generative Neural Networks

Drug Analogs from Fragment-Based Long Short-Term Memory Generative Neural Networks
复制标题

DOI:
10.1021/acs.jcim.8b00902
复制
发表时间:
2019-04-01
影响因子:
5.6
通讯作者:
Reymond, Jean-Louis
Reymond, Jean-Louis
中科院分区:
化学2区
文献类型:
--
作者:
Awale, Mahendra;Sirockin, Finton;Reymond, Jean-Louis

文献摘要

被引文献

相似文献

最近的几份报告表明,用于语法学习的长短期记忆生成神经网络(LSTM)在使用来自生物活性化合物数据库(如ChEMBL)的简化分子输入行输入系统(SMILES)进行训练时,可以有效地学习编写药物样化合物的SMILES,并且随后可以在使用具有特定生物活性特征的化合物进行转移学习时产生集中集。在这里,我们使用来自ChEMBL,DrugBank,商业可用片段或FDB-17(多达17个原子的片段数据库)的分子训练LSTM,并对单个已知药物进行迁移学习,以获得该药物的新类似物。我们发现,这种方法很容易生成数百种相关的、多样化的新药类似物,并且最适合使用约40,000种化合物的训练集,这些化合物就像商业片段一样简单。这些数据表明,基于片段的LSTM为新分子的生成提供了一种有前途的方法。
Several recent reports have shown that long short-term memory generative neural networks (LSTM) of the type used for grammar learning efficiently learn to write Simplified Molecular Input Line Entry System (SMILES) of druglike compounds when trained with SMILES from a database of bioactive compounds such as ChEMBL and can later produce focused sets upon transfer learning with compounds of specific bioactivity profiles. Here we trained an LSTM using molecules taken either from ChEMBL, DrugBank, commercially available fragments, or from FDB-17 (a database of fragments up to 17 atoms) and performed transfer learning to a single known drug to obtain new analogs of this drug. We found that this approach readily generates hundreds of relevant and diverse new drug analogs and works best with training sets of around 40,000 compounds as simple as commercial fragments. These data suggest that fragment-based LSTM offer a promising method for new molecule generation.