General melting point prediction based on a diverse compound data set and artificial neural networks

General melting point prediction based on a diverse compound data set and artificial neural networks
复制标题

DOI:
10.1021/ci0500132
复制
发表时间:
2005-05-01
影响因子:
5.6
通讯作者:
Bender, A
Bender, A
中科院分区:
化学2区
文献类型:
--
作者:
Karthikeyan, M;Glen, RC;Bender, A

文献摘要

被引文献

相似文献

我们报告了一个强大的和一般的模型预测熔点的发展。它基于4173种化合物的多样化数据集,并采用大量的2D和3D描述符来捕获分子物理化学和其他基于图形的特性。通过主成分分析法来降低非线性,同时采用全连接的前馈反向传播人工神经网络来生成模型。熔点是分子的基本物理化学性质,其由单分子性质和由于在固态中堆积而引起的分子间相互作用控制。因此,很难预测,以前只开发了明确定义和较小化合物集的熔点模型。在这里,我们得出了第一个一般模型,涵盖了一个比较大的和相关的有机化学空间的一部分。最终的模型是基于2D描述符,这被发现包含更多的相关信息比3D描述符计算。该模型的内部随机验证实现了R-2 = 0.661的相关系数,平均绝对误差为37.6 ℃。该模型内部一致,测试集的相关系数Q(2)= 0.658(平均绝对误差38.2 ℃),内部验证集的相关系数Q(2)= 0.645(平均绝对误差39.8 ℃)。对由277种化合物组成的外部药物数据集进行了额外验证。在该外部数据集上,实现了Q(2)= 0.662(平均绝对误差32.6 ℃)的相关系数,显示了模型的推广能力。与早期的模型相比,我们的模型表现出略有改善的性能,尽管覆盖的化学空间要大得多。剩余的模型误差是由于使用基于单分子的描述符未捕获的分子性质,即分子间和分子内相互作用和晶体堆积,给出了离群值的示例和原因。
We report the development of a robust and general model for the prediction of melting points. It is based on a diverse data set of 4173 compounds and employs a large number of 2D and 3D descriptors to capture molecular physicochemical and other graph-based properties. Dimensionality reduction is performed by principal component analysis, while a fully connected feed-forward back-propagation artificial neural network is employed for model generation. The melting point is a fundamental physicochemical property of a molecule that is controlled by both single-molecule properties and intermolecular interactions due to packing in the solid state. Thus, it is difficult to predict, and previously only melting point models for clearly defined and smaller compound sets have been developed. Here we derive the first general model that covers a comparatively large and relevant part of organic chemical space. The final model is based on 2D descriptors, which are found to contain more relevant information than the 3D descriptors calculated. Internal random validation of the model achieves a correlation coefficient of R-2 = 0.661 with an average absolute error of 37.6 degrees C. The model is internally consistent with a correlation coefficient of the test set of Q(2) = 0.658 (average absolute error 38.2 degrees C) and a correlation coefficient of the internal validation set of Q(2) = 0.645 (average absolute error 39.8 degrees C). Additional validation was performed on an external drug data set consisting of 277 compounds. On this external data set a correlation coefficient of Q(2) = 0.662 (average absolute error 32.6 degrees C) was achieved, showing ability of the model to generalize. Compared to an earlier model for the prediction of melting points of druglike compounds our model exhibits slightly improved performance, despite the much larger chemical space covered. The remaining model error is due to molecular properties that are not captured using single-molecule based descriptors, namely both inter- and intramolecular interactions and crystal packing, for which examples of and reasons for outliers are given.