Predicting the Thermodynamic Stability of Solids Combining Density Functional Theory and Machine Learning

Predicting the Thermodynamic Stability of Solids Combining Density Functional Theory and Machine Learning
复制标题

DOI:
10.1021/acs.chemmater.7b00156
复制
发表时间:
2017-06-27
影响因子:
8.6
通讯作者:
Marques, Miguel A. L.
Marques, Miguel A. L.
中科院分区:
材料科学2区
文献类型:
--
作者:
Schmidt, Jonathan;Shi, Jingmin;Marques, Miguel A. L.

文献摘要

被引文献

相似文献

我们执行机器学习方法的大规模基准预测固体的热力学稳定性。我们首先构建了一个数据集,其中包括大约25万立方钙钛矿系统的密度泛函理论计算。这包括所有可能的钙钛矿和反钙钛矿晶体,这些晶体可以由氢到铋的元素生成,不包括稀有气体和镧系元素。顺便说一句,这些计算已经揭示了大量的系统(大约500),它们是热力学稳定的,但在晶体结构数据库中没有。此外,其中一些相具有非常规的成分,并定义了全新的钙钛矿族。然后使用该数据集来训练和测试一系列机器学习算法,以预测到稳定凸壳的能量距离。我们特别研究了脊回归、随机森林、极度随机树(包括自适应增强)和神经网络的性能。我们发现,在230000个钙钛矿的测试集中,经过20000个样本的训练,极度随机化树给出了到凸壳距离的最小平均绝对误差(121 meV/原子)。令人惊讶的是,如果我们把组成钙钛矿的三种元素的元素周期表中的组和行作为唯一的输入特征,这个机器已经工作了。此外,我们发现预测精度在元素周期表中并不均匀,对于第一行元素和形成磁性化合物的元素来说,预测精度更差。我们的研究结果表明,机器学习可以通过限制相关化学成分的空间而不降低准确性,从而大大加快(至少5倍)高通量DFT计算。
We perform a large scale benchmark of machine learning methods for the prediction of the thermodynamic stability of solids. We start by constructing a data set that comprises density functional theory calculations of around 250000 cubic perovskite systems. This includes all possible 3 perovskite and antiperovskite crystals that can be generated with elements from hydrogen to bismuth, excluding rare gases and lanthanides. Incidentally, these calculations already reveal a large number of systems (around 500) that are thermodynamically stable but that are not present in crystal structure databases. Moreover, some of these phases have unconventional compositions and define completely new families of perovskites. This data set is then used to train and test a series of machine learning algorithms to predict the energy distance to the convex hull of stability. In particular, we study the performance of ridge regression, random forests, extremely randomized trees (including adaptive boosting), and neural networks. We find that extremely randomized trees give the smallest mean absolute error of the distance to the convex hull (121 meV/atom) in the test set of 230000 perovskites, after being trained in 20000 samples. Surprisingly, the machine already works if we give it as sole input features the group and row in the periodic table of the three elements composing the perovskite. Moreover, we find that the prediction accuracy is not uniform across the periodic table, being worse for first-row elements and elements forming magnetic compounds. Our results suggest that machine learning can be used to speed up considerably (by at least a factor of 5) high-throughput DFT calculations, by restricting the space of relevant chemical compositions without degradation of the accuracy.