Enabling deeper learning on big data for materials informatics applications.

Enabling deeper learning on big data for materials informatics applications.
复制标题

DOI:
10.1038/s41598-021-83193-1
复制
发表时间:
2021-02-19
期刊:
影响因子:
4.6
通讯作者:
Agrawal A
Agrawal A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Jha D;Gupta V;Ward L;Yang Z;Wolverton C;Foster I;Liao WK;Choudhary A;Agrawal A

文献摘要

参考文献

被引文献

相似文献

近年来,机器学习(ML)技术在材料科学中的应用引起了人们的极大关注,因为它们能够有效地从各种输入材料表示中提取数据驱动的联系,从而获得其输出属性。虽然传统机器学习技术的应用已经变得相当普遍,但更先进的深度学习(DL)技术的应用有限,主要是因为大型材料数据集相对较少。鉴于DL的潜力和优势以及大型材料数据集的可用性越来越高,为了提高模型性能而进行更深入的神经网络是很有吸引力的,但实际上,由于梯度消失问题,它会导致性能下降。在本文中,我们解决了如何对可用大材料数据的案例进行更深入学习的问题。在这里,我们提出了一个基于个体残差学习(IRNet)的通用深度学习框架,该框架由非常深的神经网络组成,可以使用任何基于矢量的材料表示作为输入,以构建准确的属性预测模型。我们发现,所提出的IRNet模型不仅可以成功地缓解消失梯度问题并实现更深入的学习,而且与普通深度神经网络和传统ML技术相比,在大数据存在的情况下,对于给定的输入材料表示,模型准确性显着提高(高达47%)。
The application of machine learning (ML) techniques in materials science has attracted significant attention in recent years, due to their impressive ability to efficiently extract data-driven linkages from various input materials representations to their output properties. While the application of traditional ML techniques has become quite ubiquitous, there have been limited applications of more advanced deep learning (DL) techniques, primarily because big materials datasets are relatively rare. Given the demonstrated potential and advantages of DL and the increasing availability of big materials datasets, it is attractive to go for deeper neural networks in a bid to boost model performance, but in reality, it leads to performance degradation due to the vanishing gradient problem. In this paper, we address the question of how to enable deeper learning for cases where big materials data is available. Here, we present a general deep learning framework based on Individual Residual learning (IRNet) composed of very deep neural networks that can work with any vector-based materials representation as input to build accurate property prediction models. We find that the proposed IRNet models can not only successfully alleviate the vanishing gradient problem and enable deeper learning, but also lead to significantly (up to 47%) better model accuracy as compared to plain deep neural networks and traditional ML techniques for a given input materials representation in the presence of big data.
DOI: 10.1063/1.4946894
发表时间: 2016-05-01
期刊: APL MATERIALS
影响因子: 6.1
作者:
Agrawal, Ankit;Choudhary, Alok
通讯作者: Choudhary, Alok
DOI: 10.1103/physrevlett.117.135502
发表时间: 2016-09-20
影响因子: 8.6
作者:
Faber, Felix A.;Lindmaa, Alexander;Armiento, Rickard
通讯作者: Armiento, Rickard
DOI: 10.1038/s41598-018-35934-y
发表时间: 2018-12-04
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Jha, Dipendra;Ward, Logan;Agrawal, Ankit
通讯作者: Agrawal, Ankit
DOI: 10.1016/j.commatsci.2012.02.002
发表时间: 2012-06-01
影响因子: 3.3
作者:
Curtarolo, Stefano;Setyawan, Wahyu;Levy, Ohad
通讯作者: Levy, Ohad
DOI: 10.1063/1.4812323
发表时间: 2013-07-01
期刊: APL MATERIALS
影响因子: 6.1
作者:
Jain, Anubhav;Shyue Ping Ong;Persson, Kristin A.
通讯作者: Persson, Kristin A.