Balancing data for generalizable machine learning to predict glass-forming ability of ternary alloys

Balancing data for generalizable machine learning to predict glass-forming ability of ternary alloys
复制标题

DOI:
10.1016/j.scriptamat.2021.114366
复制
发表时间:
2022
期刊:
影响因子:
6
通讯作者:
Yi Yao;Timothy Sullivan;Feng Yan;Jiaqi Gong;Lin Li
Yi Yao;Timothy Sullivan;Feng Yan;Jiaqi Gong;Lin Li
中科院分区:
材料科学1区
文献类型:
--
作者:
Yi Yao;Timothy Sullivan;Feng Yan;Jiaqi Gong;Lin Li

文献摘要

相似文献

机器学习随着数据驱动材料科学的出现而蓬勃发展。然而,现有研究工作获得的材料数据集存在严重的不平衡问题。本文研究了三元合金体系玻璃形成能力的数据不平衡,包括丰富的、低保真度的高通量数据和稀疏的、高保真度的传统实验数据。我们演示了一种处理数据不平衡的新方法,并在原始数据集与平衡数据集上训练了人工神经网络 (ANN) 模型。在平衡数据集上训练的 ANN 模型解决了在原始数据集上训练的模型所遭受的过度拟合问题。更重要的是,数据平衡模型预测新合金系统的普遍性得到了提高,留一合金系统验证证明了这一点。我们的工作强调了处理材料数据集中的数据不平衡的重要性,以解决机器学习模型的过度拟合问题,并进一步增强预测新材料系统特性的通用性。
Machine Learning has thrived on the emergence of data-driven materials science. However, the materials datasets acquired at existing research efforts have significant imbalance issues. This paper investigated the data imbalance for the glass-forming ability of ternary alloy systems, which consists of abundant, low-fidelity high-throughput data, and sparse, high-fidelity traditional experimental data. We demonstrated a new method to handle the data imbalance and trained artificial neural network (ANN) models on the original vs. balanced datasets. The ANN model trained on the balanced dataset solved the overfitting issue suffered by the model trained on the original dataset. More importantly, the generalizability in predicting the new alloy system was improved in the data-balanced model, evidenced by the leave-one-alloy-system-out validation. Our work highlights the importance of handling data imbalance in material datasets to solve the overfitting issues of machine learning models and further enhance generalizability in predicting the characteristics of the new material systems.