A Fragmentation-Based Graph Embedding Framework for QM/ML

A Fragmentation-Based Graph Embedding Framework for QM/ML
复制标题

用于 QM/ML 的基于碎片的图嵌入框架

DOI:
10.1021/acs.jpca.1c06152
复制
发表时间:
2021
期刊:
The Journal of Physical Chemistry A
影响因子:
--
通讯作者:
Raghavachari, Krishnan
Raghavachari, Krishnan
中科院分区:
--
文献类型:
--
作者:
Collins, Eric M.;Raghavachari, Krishnan

文献摘要

相似文献

我们为 QM/ML 方法引入了一种新的基于碎片的分子表示框架“FragGraph”,涉及将碎片指纹嵌入到分子图上。我们的模型专为增量机器学习(Δ-ML)而设计,其中心目标是纠正 DFT 等近似方法的缺陷,以实现高精度。我们的框架基于碎片、错误消除和最先进的深度学习架构的思想的明智组合。从广义上讲,我们通过将预先构建的固有优势融入错误消除方法(例如广义的基于连接的层次结构)中,开发了用于分子机器学习的通用图网络框架。更具体地说,我们通过使用片段分子指纹编码的基于片段的属性图表示来开发 QM/ML 表示。我们的表示的实用性通过图网络指纹编码器得到证明,其中全局指纹是通过片段指纹的局部邻域的消息传递生成的,有效地增强了标准指纹以包括内置的分子图结构。在 130k-GDB9 数据集上,我们的方法预测样本外平均绝对误差与目标 G4(MP2) 计算的能量相比显着低于 1 kJ/mol,可与当前的深度学习方法相媲美,并减少计算规模。
We introduce a new fragmentation-based molecular representation framework “FragGraph” for QM/ML methods involving embedding fragment-wise fingerprints onto molecular graphs. Our model is specifically designed for delta machine learning (Δ-ML) with the central goal of correcting the deficiencies of approximate methods such as DFT to achieve high accuracy. Our framework is based on a judicious combination of ideas from fragmentation, error cancellation, and a state-of-the-art deep learning architecture. Broadly, we develop a general graph-network framework for molecular machine learning by incorporating the inherent advantages prebuilt into error cancellation methods such as the generalized Connectivity-Based Hierarchy. More specifically, we develop a QM/ML representation through a fragmentation-based attributed graph representation encoded with fragment-wise molecular fingerprints. The utility of our representation is demonstrated through a graph network fingerprint encoder in which a global fingerprint is generated through message passing of local neighborhoods of fragment-wise fingerprints, effectively augmenting standard fingerprints to also include the inbuilt molecular graph structure. On the 130k-GDB9 dataset, our method predicts an out-of-sample mean absolute error significantly lower than 1 kJ/mol compared to target G4(MP2) calculated energies, rivaling current deep learning methods with reduced computational scaling.