FAME: Fragment-based Conditional Molecular Generation for Phenotypic Drug Discovery.
FAME: Fragment-based Conditional Molecular Generation for Phenotypic Drug Discovery.
复制标题
DOI:
10.1137/1.9781611977172.81
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
中科院分区:
文献类型:
--
作者:
De novo molecular design is a key challenge in drug discovery due to the complexity of chemical space. With the availability of molecular datasets and advances in machine learning, many deep generative models are proposed for generating novel molecules with desired properties. However, most of the existing models focus only on molecular distribution learning and target-based molecular design, thereby hindering their potentials in real-world applications. In drug discovery, phenotypic molecular design has advantages over target-based molecular design, especially in first-in-class drug discovery. In this work, we propose the first deep graph generative model (FAME) targeting phenotypic molecular design, in particular gene expression-based molecular design. FAME leverages a conditional variational autoencoder framework to learn the conditional distribution generating molecules from gene expression profiles. However, this distribution is difficult to learn due to the complexity of the molecular space and the noisy phenomenon in gene expression data. To tackle these issues, a gene expression denoising (GED) model that employs contrastive objective function is first proposed to reduce noise from gene expression data. FAME is then designed to treat molecules as the sequences of fragments and learn to generate these fragments in autoregressive manner. By leveraging this fragment-based generation strategy and the denoised gene expression profiles, FAME can generate novel molecules with a high validity rate and desired biological activity. The experimental results show that FAME outperforms existing methods including both SMILES-based and graph-based deep generative models for phenotypic molecular design. Furthermore, the effective mechanism for reducing noise in gene expression data proposed in our study can be applied to omics data modeling in general for facilitating phenotypic drug discovery.
登录
查看更多内容
影响因子:
18.2
作者:
Segler MHS;Kogej T;Tyrchan C;Waller MP
通讯作者:
Waller MP
影响因子:
18.2
作者:
Gómez-Bombarelli R;Wei JN;Duvenaud D;Hernández-Lobato JM;Sánchez-Lengeling B;Sheberla D;Aguilera-Iparraguirre J;Hirzel TD;Adams RP;Aspuru-Guzik A
通讯作者:
Aspuru-Guzik A
影响因子:
8.6
作者:
Sun J;Jeliazkova N;Chupakin V;Golib-Dzib JF;Engkvist O;Carlsson L;Wegner J;Ceulemans H;Georgiev I;Jeliazkov V;Kochev N;Ashby TJ;Chen H
通讯作者:
Chen H
影响因子:
14.9
作者:
Gaulton A;Hersey A;Nowotka M;Bento AP;Chambers J;Mendez D;Mutowo P;Atkinson F;Bellis LJ;Cibrián-Uhalte E;Davies M;Dedman N;Karlsson A;Magariños MP;Overington JP;Papadatos G;Smit I;Leach AR
通讯作者:
Leach AR
影响因子:
16.6
作者:
Mendez-Lucio, Oscar;Baillif, Benoit;Wichard, Joerg
通讯作者:
Wichard, Joerg