Synthetic pre-training for neural-network interatomic potentials

Synthetic pre-training for neural-network interatomic potentials
复制标题

DOI:
10.1088/2632-2153/ad1626
复制
发表时间:
2023-07
期刊:
Machine Learning: Science and Technology
影响因子:
--
通讯作者:
John L A Gardner;Kathryn T. Baker;Volker L. Deringer
John L A Gardner;Kathryn T. Baker;Volker L. Deringer
中科院分区:
其他
文献类型:
--
作者:
John L A Gardner;Kathryn T. Baker;Volker L. Deringer

文献摘要

被引文献

相似文献

基于机器学习(ML)的原子间势已经改变了原子材料建模领域。然而,机器学习的潜力很大程度上取决于训练它们的量子力学参考数据的质量和数量,因此开发数据集和训练管道正成为一个日益重要的挑战。利用在机器学习研究的其他领域中常见的“合成”(人工)数据的想法,我们在这里展示了合成原子数据,它们本身是用现有的机器学习潜力大规模获得的,构成了神经网络(NN)原子间势模型的有用的预训练任务。一旦使用大型合成数据集进行预训练,这些模型就可以在更小的量子力学数据集上进行微调,从而提高计算实践中的数值准确性和稳定性。我们证明了碳的一系列等变图-神经网络电位的可行性,并进行了初步实验来测试该方法的局限性。
Machine learning (ML) based interatomic potentials have transformed the field of atomistic materials modelling. However, ML potentials depend critically on the quality and quantity of quantum-mechanical reference data with which they are trained, and therefore developing datasets and training pipelines is becoming an increasingly central challenge. Leveraging the idea of ‘synthetic’ (artificial) data that is common in other areas of ML research, we here show that synthetic atomistic data, themselves obtained at scale with an existing ML potential, constitute a useful pre-training task for neural-network (NN) interatomic potential models. Once pre-trained with a large synthetic dataset, these models can be fine-tuned on a much smaller, quantum-mechanical one, improving numerical accuracy and stability in computational practice. We demonstrate feasibility for a series of equivariant graph-NN potentials for carbon, and we carry out initial experiments to test the limits of the approach.