The Role of Structural Representation in the Performance of a Deep Neural Network for X-ray Spectroscopy

The Role of Structural Representation in the Performance of a Deep Neural Network for X-ray Spectroscopy
复制标题

DOI:
10.3390/molecules25112715
复制
发表时间:
2020-06
期刊:
影响因子:
4.6
通讯作者:
Marwah M M Madkhali-Marwah-M-M-Madkhali-93008847;C. Rankine;T. Penfold
Marwah M M Madkhali-Marwah-M-M-Madkhali-93008847;C. Rankine;T. Penfold
中科院分区:
化学2区
文献类型:
--
作者:
Marwah M M Madkhali-Marwah-M-M-Madkhali-93008847;C. Rankine;T. Penfold

文献摘要

被引文献

相似文献

开发用于预测分子特性的深度神经网络 (DNN) 时的一个重要考虑因素是化学空间的表示。在此,我们探讨了表征对 DNN 性能的影响,该 DNN 旨在预测 Fe K 边缘 X 射线吸收近边缘结构 (XANES) 光谱,并解决了以下问题:任意 Fe 吸收位点周围局部环境的表征选择有多重要?使用两种流行的化学空间表示——库仑矩阵 (CM) 和对分布/径向分布曲线 (RDC)——我们研究了表示选择对 DNN 性能的影响。虽然 CM 和 RDC 特征化是明显稳健的描述符,但在使用 RDC 特征化时,可以在目标和估计的 XANES 光谱之间获得更小的均方误差 (MSE),并且 a) 更快 b) 使用更少的数据样本收敛到此状态。这有利于我们的 DNN 未来扩展到其他 X 射线吸收边缘,以及重新优化我们的 DNN 以重现更高水平理论的结果。在后一种情况下,数据集大小将受到基础理论计算的资源密集型性质的更大限制。
An important consideration when developing a deep neural network (DNN) for the prediction of molecular properties is the representation of the chemical space. Herein we explore the effect of the representation on the performance of our DNN engineered to predict Fe K-edge X-ray absorption near-edge structure (XANES) spectra, and address the question: How important is the choice of representation for the local environment around an arbitrary Fe absorption site? Using two popular representations of chemical space—the Coulomb matrix (CM) and pair-distribution/radial distribution curve (RDC)—we investigate the effect that the choice of representation has on the performance of our DNN. While CM and RDC featurisation are demonstrably robust descriptors, it is possible to obtain a smaller mean squared error (MSE) between the target and estimated XANES spectra when using RDC featurisation, and converge to this state a) faster and b) using fewer data samples. This is advantageous for future extension of our DNN to other X-ray absorption edges, and for reoptimisation of our DNN to reproduce results from higher levels of theory. In the latter case, dataset sizes will be limited more strongly by the resource-intensive nature of the underlying theoretical calculations.