Comparing predictive ability of QSAR/QSPR models using 2D and 3D molecular representations

Comparing predictive ability of QSAR/QSPR models using 2D and 3D molecular representations
复制标题

DOI:
10.1007/s10822-020-00361-7
复制
发表时间:
2021-01
影响因子:
3.5
通讯作者:
A. Sato;Tomoyuki Miyao;Swarit Jasial;K. Funatsu
A. Sato;Tomoyuki Miyao;Swarit Jasial;K. Funatsu
中科院分区:
生物学3区
文献类型:
--
作者:
A. Sato;Tomoyuki Miyao;Swarit Jasial;K. Funatsu

文献摘要

相似文献

定量结构-活性关系(QSAR)和定量结构-性质关系(QSPR)模型是基于化学结构与活性(性质)值之间的数值关系来预测生物活性和分子性质的。分子表征在QSAR/QSPR分析中具有重要意义。分子结构的拓扑信息通常用于此目的(2D表示)。然而,构象信息似乎很重要,因为分子是在三维空间中。作为适用于不同化合物的三维分子表示,先前已经提出了测试分子和一组参考分子之间的相似性。发现这种3D表示对于活性化合物的早期富集的虚拟筛选是有效的。在本研究中,我们将三维表示引入QSAR/QSPR建模(回归任务)。此外,我们研究了3D表示在训练数据集的多样性方面相对于2D的优点。对于基于量子力学的性质的预测任务,3D表示上级2D。对于预测小分子对特定生物靶标的活性,无论训练数据集的多样性如何,使用两种类型的表示在性能差异中没有观察到一致的趋势。
Quantitative structure–activity relationship (QSAR) and quantitative structure–property relationship (QSPR) models predict biological activity and molecular property based on the numerical relationship between chemical structures and activity (property) values. Molecular representations are of importance in QSAR/QSPR analysis. Topological information of molecular structures is usually utilized (2D representations) for this purpose. However, conformational information seems important because molecules are in the three-dimensional space. As a three-dimensional molecular representation applicable to diverse compounds, similarity between a test molecule and a set of reference molecules has been previously proposed. This 3D representation was found to be effective on virtual screening for early enrichment of active compounds. In this study, we introduced the 3D representation into QSAR/QSPR modeling (regression tasks). Furthermore, we investigated relative merits of 3D representations over 2D in terms of the diversity of training data sets. For the prediction task of quantum mechanics-based properties, the 3D representations were superior to 2D. For predicting activity of small molecules against specific biological targets, no consistent trend was observed in the difference of performance using the two types of representations, irrespective of the diversity of training data sets.