VLA-SMILES: Variable-Length-Array SMILES Descriptors in Neural Network-Based QSAR Modeling

VLA-SMILES: Variable-Length-Array SMILES Descriptors in Neural Network-Based QSAR Modeling
复制标题

DOI:
10.3390/make4030034
复制
发表时间:
2022-08
期刊:
Mach. Learn. Knowl. Extr.
影响因子:
--
通讯作者:
Antonina L. Nazarova;A. Nakano
Antonina L. Nazarova;A. Nakano
中科院分区:
其他
文献类型:
--
作者:
Antonina L. Nazarova;A. Nakano

文献摘要

相似文献

机器学习是数据驱动研究的一个里程碑,包括材料信息学、机器人技术和计算机辅助药物发现。随着虚拟化学空间和合成化学空间的不断扩大,人们需要高效、稳健的定量构效关系(QSAR)方法来揭示具有所需性质的分子。在这里,我们提出了基于可变长度数组SMILES(VLA-SMILES)的结构描述符,扩展了广泛用于机器学习的传统SMILES描述符。这种结构表示扩展了数字编码的SMILES家族,特别是二进制SMILES,以加快发现具有高预测能力的新深度学习QSAR模型。VLA-SMILES描述符被证明可以加速基于多层感知器(MLP)的QSAR模型的训练,优化反向传播(ATransformedBP),弹性传播(iRPROP)和Adam优化学习算法具有合理的训练测试分裂,同时提高对计算密集型二进制SMILES表示格式的预测能力。所有测试的MLP在相同的基于长度阵列的SMILES描述符下显示出相似的预测能力和收敛速度的训练相结合的考虑学习程序。基于结构描述符相似性度量的Kennard-Stone训练测试分裂验证被发现比基于整个VLA-SMILES特征QSAR集的生物活性值度量的活性排序分区更有效。通过QSAR参数模型验证方法对基于VLA-SMILES的MLP模型的稳健性和预测能力进行了评价。此外,基于F2,n−2 -标准的真实的和观察到的活性之间线性回归的统计H 0假设检验方法用于VLA-SMILES特征QSAR-MLP的可预测性估计(n为检验集的体积)。这两种方法的QSAR参数模型验证和统计假设检验被发现相关时,用于定量评价的可预测性的设计QSAR模型与VLA-SMILES描述符。
Machine learning represents a milestone in data-driven research, including material informatics, robotics, and computer-aided drug discovery. With the continuously growing virtual and synthetically available chemical space, efficient and robust quantitative structure–activity relationship (QSAR) methods are required to uncover molecules with desired properties. Herein, we propose variable-length-array SMILES-based (VLA-SMILES) structural descriptors that expand conventional SMILES descriptors widely used in machine learning. This structural representation extends the family of numerically coded SMILES, particularly binary SMILES, to expedite the discovery of new deep learning QSAR models with high predictive ability. VLA-SMILES descriptors were shown to speed up the training of QSAR models based on multilayer perceptron (MLP) with optimized backpropagation (ATransformedBP), resilient propagation (iRPROP‒), and Adam optimization learning algorithms featuring rational train–test splitting, while improving the predictive ability toward the more compute-intensive binary SMILES representation format. All the tested MLPs under the same length-array-based SMILES descriptors showed similar predictive ability and convergence rate of training in combination with the considered learning procedures. Validation with the Kennard–Stone train–test splitting based on the structural descriptor similarity metrics was found more effective than the partitioning with the ranking by activity based on biological activity values metrics for the entire set of VLA-SMILES featured QSAR. Robustness and the predictive ability of MLP models based on VLA-SMILES were assessed via the method of QSAR parametric model validation. In addition, the method of the statistical H0 hypothesis testing of the linear regression between real and observed activities based on the F2,n−2 -criteria was used for predictability estimation among VLA-SMILES featured QSAR-MLPs (with n being the volume of the testing set). Both approaches of QSAR parametric model validation and statistical hypothesis testing were found to correlate when used for the quantitative evaluation of predictabilities of the designed QSAR models with VLA-SMILES descriptors.