Structure classification and melting temperature prediction in octet AB solids via machine learning

Structure classification and melting temperature prediction in octet AB solids via machine learning
复制标题

DOI:
10.1103/physrevb.91.214302
复制
发表时间:
2015-06
期刊:
影响因子:
3.7
通讯作者:
G. Pilania;J. Gubernatis;T. Lookman
G. Pilania;J. Gubernatis;T. Lookman
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
G. Pilania;J. Gubernatis;T. Lookman

文献摘要

被引文献

相似文献

机器学习方法在凝聚态物理和材料科学中越来越多地用于分类晶体结构和预测材料性质。然而,这些方法对于给定问题的可靠性,特别是在没有大量数据集的情况下,还没有得到很好的研究。通过解决分类晶体结构和预测AB固体八元子集的熔化温度的任务,我们进行了这样的研究,并发现了在相对较小的数据集上使用机器学习方法的潜在问题。然而,与此同时,我们可以重申这些方法对这些任务的潜在力量。特别是,我们发现了一个重要的新材料特征,即超额玻恩有效电荷,这大大提高了我们定义的分类问题预测的准确性。这一发现使我们提出了这些固体中离子性和共价程度的新尺度。更具体地说,我们将一组75个八隅体固体的晶体结构划分为离子和共价键,从而执行了二元分类任务。我们发现,使用圣约翰和布洛赫几十年前提出的标准指数(r σ, r π),可以使分类的平均成功率达到92%。我们发现,仅使用r σ和A原子的超额玻恩有效电荷∆Z A,平均成功率为97%,但我们也发现,这些平均值的变化相对较大,这取决于某些机器学习方法的使用方式,而标准偏差并不是衡量我们可以放置在任何平均值中的置信度的适当指标。相反,我们计算并以95%的置信度报告,传统分类对预测的准确率在区间[89%,95%],新分类对的准确率在区间[96%,99%]。对于熔化温度的预测,我们的数据集的大小是46。我们估计结果模型的均方根误差为数据平均熔化温度的11%,但我们注意到,如果该预测误差的精度本身是测量的,我们估计的拟合误差本身的均方根误差为50%。简而言之,我们所说明的是,分类和回归预测可以根据细节发生显著变化
Machine learning methods are being increasingly used in condensed matter physics and materials science to classify crystals structures and predict material properties. However, the reliability of these methods for a given problem, especially when large data sets are unavailable, has not been well studied. By addressing the tasks of classifying crystal structure and predicting melting temperatures of the octet subset of AB solids, we performed such a study and found potential problems with using machine learning methods on relatively small data sets. At the same time, however, we can reaffirm the potential power of such methods for these tasks. In particular, we uncovered an important new material feature, the excess Born effective charge, that significantly increased the accuracy of the predictions for the classification problem we defined. This discovery leads us to propose a new scale for the degree of ionicity and covalency in these solids. More specifically, we partitioned the crystal structures of a set of 75 octet solids into those that are ionic and covalent bonded and thus performed a binary classification task. We found that using the standard indices ( r σ , r π ), suggested by St. John and Bloch several decades ago, enabled an average success in classification of 92%. We found that using just r σ and the excess Born effective charge ∆ Z A of the A atom enabled an average success of 97%, but we also found relatively large variations about these averages that were dependent on how certain machine learning methods were used and for which a standard deviation was not a proper measure of the degree of confidence we can place in either average. Instead, we calculated and report with 95% confidence that the traditional classification pair predicts an accuracy in the interval [89% , 95%] and the accuracy of the new pair lies in the interval [96% , 99%]. For melting temperature predictions, the size of our data set was 46. We estimate the root-mean-squared error of our resulting model to be 11% of the mean melting temperature of the data, but we note that if the accuracy of this predicted error is itself measured, our estimated fitting error itself has root-mean-square error of 50%. In short, what we illustrate is that classification and regression predictions can vary significantly, depending on the details