FAME: Fragment-based Conditional Molecular Generation for Phenotypic Drug Discovery.

FAME: Fragment-based Conditional Molecular Generation for Phenotypic Drug Discovery.
复制标题

DOI:
10.1137/1.9781611977172.81
复制
发表时间:
2022
期刊:
Proceedings of the ... SIAM International Conference on Data Mining. SIAM International Conference on Data Mining
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

由于化学空间的复杂性,从头分子设计是药物发现的一个关键挑战。随着分子数据集的可用性和机器学习的进步,人们提出了许多深度生成模型来生成具有所需特性的新型分子。然而,大多数现有模型仅关注分子分布学习和基于目标的分子设计,从而阻碍了它们在实际应用中的潜力。在药物发现中,表型分子设计比基于靶标的分子设计具有优势,尤其是在一流药物发现中。在这项工作中,我们提出了第一个针对表型分子设计,特别是基于基因表达的分子设计的深度图生成模型(FAME)。 FAME 利用条件变分自动编码器框架来学习从基因表达谱生成分子的条件分布。然而,由于分子空间的复杂性和基因表达数据中的噪声现象,这种分布很难学习。为了解决这些问题,首先提出了一种采用对比目标函数的基因表达去噪(GED)模型来减少基因表达数据中的噪声。然后,FAME 被设计为将分子视为片段序列,并学习以自回归方式生成这些片段。通过利用这种基于片段的生成策略和去噪基因表达谱,FAME 可以生成具有高有效性和所需生物活性的新型分子。实验结果表明,FAME 优于现有方法,包括用于表型分子设计的基于 SMILES 和基于图的深度生成模型。此外,我们的研究中提出的减少基因表达数据噪音的有效机制可以应用于一般的组学数据建模,以促进表型药物的发现。
De novo molecular design is a key challenge in drug discovery due to the complexity of chemical space. With the availability of molecular datasets and advances in machine learning, many deep generative models are proposed for generating novel molecules with desired properties. However, most of the existing models focus only on molecular distribution learning and target-based molecular design, thereby hindering their potentials in real-world applications. In drug discovery, phenotypic molecular design has advantages over target-based molecular design, especially in first-in-class drug discovery. In this work, we propose the first deep graph generative model (FAME) targeting phenotypic molecular design, in particular gene expression-based molecular design. FAME leverages a conditional variational autoencoder framework to learn the conditional distribution generating molecules from gene expression profiles. However, this distribution is difficult to learn due to the complexity of the molecular space and the noisy phenomenon in gene expression data. To tackle these issues, a gene expression denoising (GED) model that employs contrastive objective function is first proposed to reduce noise from gene expression data. FAME is then designed to treat molecules as the sequences of fragments and learn to generate these fragments in autoregressive manner. By leveraging this fragment-based generation strategy and the denoised gene expression profiles, FAME can generate novel molecules with a high validity rate and desired biological activity. The experimental results show that FAME outperforms existing methods including both SMILES-based and graph-based deep generative models for phenotypic molecular design. Furthermore, the effective mechanism for reducing noise in gene expression data proposed in our study can be applied to omics data modeling in general for facilitating phenotypic drug discovery.
通过复发性神经网络生成聚焦的分子库来发现药物。
DOI: 10.1021/acscentsci.7b00512
发表时间: 2018-01-24
影响因子: 18.2
作者:
Segler MHS;Kogej T;Tyrchan C;Waller MP
通讯作者: Waller MP
DOI: 10.1021/acscentsci.7b00572
发表时间: 2018-02-28
影响因子: 18.2
作者:
Gómez-Bombarelli R;Wei JN;Duvenaud D;Hernández-Lobato JM;Sánchez-Lengeling B;Sheberla D;Aguilera-Iparraguirre J;Hirzel TD;Adams RP;Aspuru-Guzik A
通讯作者: Aspuru-Guzik A
DOI: 10.1186/s13321-017-0203-5
发表时间: 2017
影响因子: 8.6
作者:
Sun J;Jeliazkova N;Chupakin V;Golib-Dzib JF;Engkvist O;Carlsson L;Wegner J;Ceulemans H;Georgiev I;Jeliazkov V;Kochev N;Ashby TJ;Chen H
通讯作者: Chen H
DOI: 10.1093/nar/gkw1074
发表时间: 2017-01-04
影响因子: 14.9
作者:
Gaulton A;Hersey A;Nowotka M;Bento AP;Chambers J;Mendez D;Mutowo P;Atkinson F;Bellis LJ;Cibrián-Uhalte E;Davies M;Dedman N;Karlsson A;Magariños MP;Overington JP;Papadatos G;Smit I;Leach AR
通讯作者: Leach AR
DOI: 10.1038/s41467-019-13807-w
发表时间: 2020-01-03
影响因子: 16.6
作者:
Mendez-Lucio, Oscar;Baillif, Benoit;Wichard, Joerg
通讯作者: Wichard, Joerg