Inverse QSPR/QSAR Analysis for Chemical Structure Generation (from y to x)

Inverse QSPR/QSAR Analysis for Chemical Structure Generation (from y to x)
复制标题

DOI:
10.1021/acs.jcim.5b00628
复制
发表时间:
2016-02-01
影响因子:
5.6
通讯作者:
Funatsu, Kimito
Funatsu, Kimito
中科院分区:
化学2区
文献类型:
--
作者:
Miyao, Tomoyuki;Kaneko, Hiromasa;Funatsu, Kimito

文献摘要

被引文献

相似文献

从目标变量(y)的值中检索描述符信息(x信息)是逆定量结构属性关系(逆QSPR)分析中的一个基本问题,但由于原像函数的复杂性而具有挑战性。因此,我们建议使用聚类多元线性回归(cMLR)模型作为逆 QSPR 分析的 QSPR 模型。 x 信息通过结合 cMLR 和用混合高斯 (GMM) 建模的先验分布来获取概率密度函数。进行了三个案例研究来证明 cMLR 潜力的各个方面。结果发现,cMLR 的预测能力优于 MLR,特别是对于具有非线性的数据。此外,事实证明,可以考虑适用性域,因为后验分布继承了先验分布的特征(即训练数据特征)并表示具有所需属性的可能性。最后,利用 GMM/cMLR 进行了一系列反分析,旨在生成具有特定水溶性的从头结构。
Retrieving descriptor information (x information) from a value of an objective variable (y) is a fundamental problem in inverse quantitative structure property relationship (inverse-QSPR) analysis but challenging because of the complexity of the preimage function. Herewith, we propose using a cluster-wise multiple linear regression (cMLR) model as a QSPR model for inverse-QSPR analysis. x information is acquired as a probability density function by combining cMLR and the prior distribution modeled with a mixture of Gaussians (GMMs). Three case studies were conducted to demonstrate various aspects of the potential of cMLR. It was found that the predictive power of cMLR was superior to that of MLR, especially for data with nonlinearity. Moreover, it turned out that the applicability domain could be considered since the posterior distribution inherits the prior distribution's feature (i.e., training data feature) and represents the possibility of having the desired property. Finally, a series of inverse analyses with the GMMs/cMLR was demonstrated with the aim to generate de novo structures having specific aqueous solubility.