Interpretable Graph-Network-Based Machine Learning Models via Molecular Fragmentation

Interpretable Graph-Network-Based Machine Learning Models via Molecular Fragmentation
复制标题

通过分子碎片可解释的基于图网络的机器学习模型

DOI:
10.1021/acs.jctc.2c01308
复制
发表时间:
2023
影响因子:
5.5
通讯作者:
Raghavachari, Krishnan
Raghavachari, Krishnan
中科院分区:
化学1区
文献类型:
--
作者:
Collins, Eric M.;Raghavachari, Krishnan

文献摘要

相似文献

化学家长期以来一直受益于理解和解释计算模型预测的能力。随着目前向更复杂的深度学习模型的转变,在许多情况下,这种实用性已经丧失。在这项工作中,我们扩展了我们以前在计算热化学方面的工作,并提出了一个可解释的图形网络,FragGraph(节点),它提供了分解预测到片段的贡献。我们证明了我们的模型在预测校正密度泛函理论(DFT)计算的原子化能量使用Δ学习的有用性。我们的模型预测G4(MP2)质量的热化学与精度<1 kJ mol-1的GDB 9数据集。除了我们的预测的高精度,我们观察到的趋势,定量描述的缺陷B3 LYP的碎片校正。节点预测显著优于我们之前从全局状态向量进行的模型预测。当我们通过在更多样化的测试集上进行预测来探索一般性时,这种效果最为明显,这表明节点预测对将机器学习模型扩展到更大的分子不太敏感。
Chemists have long benefitted from the ability to understand and interpret the predictions of computational models. With the current shift to more complex deep learning models, in many situations that utility is lost. In this work, we expand on our previously work on computational thermochemistry and propose an interpretable graph network, FragGraph(nodes), that provides decomposed predictions into fragment-wise contributions. We demonstrate the usefulness of our model in predicting a correction to density functional theory (DFT)-calculated atomization energies using Δ-learning. Our model predicts G4(MP2)-quality thermochemistry with an accuracy of <1 kJ mol–1for the GDB9 dataset. Besides the high accuracy of our predictions, we observe trends in the fragment corrections which quantitatively describe the deficiencies of B3LYP. Node-wise predictions significantly outperform our previous model predictions from a global state vector. This effect is most pronounced as we explore the generality by predicting on more diverse test sets indicating node-wise predictions are less sensitive to extending machine learning models to larger molecules.