Synthetic data enable experiments in atomistic machine learning

Synthetic data enable experiments in atomistic machine learning
复制标题

合成数据使原子机器学习实验成为可能

DOI:
10.1039/d2dd00137c
复制
发表时间:
2023-06-12
期刊:
DIGITAL DISCOVERY
影响因子:
--
通讯作者:
Deringer, Volker L.
Deringer, Volker L.
中科院分区:
其他
文献类型:
--
作者:
Gardner, John L. A.;Beaulieu, Zoe Faure;Deringer, Volker L.

文献摘要

被引文献

相似文献

机器学习模型越来越多地用于预测化学系统中原子的特性。在为这项任务开发描述符和回归框架方面已经取得了重大进展,通常是从(相对)较小的量子力学参考数据集开始。此类更大的数据集正在变得可用,但生成成本仍然很高。在这里,我们演示了大型数据集的使用,该数据集是通过现有 ML 势模型中的每个原子能量“合成”标记的。与量子力学基本事实相比,这个过程的成本低廉,使我们能够生成数百万个数据点,进而能够对从小数据到大数据范围的原子机器学习模型进行快速实验。这种方法使我们能够深入比较回归框架,并探索基于学习表示的可视化。我们还表明,学习合成数据标签可以成为后续对小数据集进行微调的有用的预训练任务。将来,我们期望我们的开源数据集和类似的数据集将有助于在化学数据丰富的情况下快速探索深度学习模型。我们引入了使用快速机器学习模型生成的原子结构和能量的大型“合成”数据集,并证明了其对于化学中监督和无监督 ML 任务的有用性。
Machine-learning models are increasingly used to predict properties of atoms in chemical systems. There have been major advances in developing descriptors and regression frameworks for this task, typically starting from (relatively) small sets of quantum-mechanical reference data. Larger datasets of this kind are becoming available, but remain expensive to generate. Here we demonstrate the use of a large dataset that we have "synthetically" labelled with per-atom energies from an existing ML potential model. The cheapness of this process, compared to the quantum-mechanical ground truth, allows us to generate millions of datapoints, in turn enabling rapid experimentation with atomistic ML models from the small- to the large-data regime. This approach allows us here to compare regression frameworks in depth, and to explore visualisation based on learned representations. We also show that learning synthetic data labels can be a useful pre-training task for subsequent fine-tuning on small datasets. In the future, we expect that our open-sourced dataset, and similar ones, will be useful in rapidly exploring deep-learning models in the limit of abundant chemical data.We introduce a large "synthetic" dataset of atomistic structures and energies, generated using a fast machine-learning model, and we demonstrate its usefulness for supervised and unsupervised ML tasks in chemistry.