Benchmark AFLOW Data Sets for Machine Learning

Benchmark AFLOW Data Sets for Machine Learning
复制标题

DOI:
10.1007/s40192-020-00174-4
复制
发表时间:
2020-06-01
影响因子:
3.3
通讯作者:
Sparks, Taylor D.
Sparks, Taylor D.
中科院分区:
材料科学3区
文献类型:
--
作者:
Clement, Conrad L.;Kauwe, Steven K.;Sparks, Taylor D.

文献摘要

被引文献

相似文献

材料信息学越来越多地寻找利用机器学习算法的方法。决策树、集成方法、支持向量机和各种神经网络架构等技术用于预测可能的材料特性和属性值。辅以实验室合成,机器学习在化合物发现和表征中的应用代表了材料信息学中最有前途的研究方向之一。目前这种趋势的一个缺点是缺乏用于训练、验证和测试模型有效性的标准化材料数据集。应用机器学习研究依赖于基准数据来理解其结果。固定的、预定的数据集可以进行严格的模型评估和比较。不引用基准的机器学习出版物通常很难结合上下文和重现。在这篇数据描述符文章中,我们展示了取自 AFLOW 数据库的不同材料属性的数据集集合。我们描述它们、生成它们的过程以及它们作为潜在基准的用途。我们提供了一个压缩的 ZIP 文件,其中包含数据集和相关 Python 代码的 GitHub 存储库。最后,我们讨论了未来工作整合数据集和创建类似基准集合的机会。
Materials informatics is increasingly finding ways to exploit machine learning algorithms. Techniques such as decision trees, ensemble methods, support vector machines, and a variety of neural network architectures are used to predict likely material characteristics and property values. Supplemented with laboratory synthesis, applications of machine learning to compound discovery and characterization represent one of the most promising research directions in materials informatics. A shortcoming of this trend, in its current form, is a lack of standardized materials data sets on which to train, validate, and test model effectiveness. Applied machine learning research depends on benchmark data to make sense of its results. Fixed, predetermined data sets allow for rigorous model assessment and comparison. Machine learning publications that do not refer to benchmarks are often hard to contextualize and reproduce. In this data descriptor article, we present a collection of data sets of different material properties taken from the AFLOW database. We describe them, the procedures that generated them, and their use as potential benchmarks. We provide a compressed ZIP file containing the data sets and a GitHub repository of associated Python code. Finally, we discuss opportunities for future work incorporating the data sets and creating similar benchmark collections.