Prediction of the coefficient of linear thermal expansion for the amorphous homopolymers based on chemical structure using machine learning

Prediction of the coefficient of linear thermal expansion for the amorphous homopolymers based on chemical structure using machine learning
复制标题

使用机器学习根据化学结构预测非晶均聚物的线性热膨胀系数

DOI:
10.1080/27660400.2021.1993729
复制
发表时间:
2021
期刊:
Science and Technology of Advanced Materials: Methods
影响因子:
--
通讯作者:
Nakata Ayako
Nakata Ayako
中科院分区:
--
文献类型:
--
作者:
Gracheva Ekaterina;Lambard Guillaume;Samitsu Sadaki;Sodeyama Keitaro;Nakata Ayako

文献摘要

参考文献

相似文献

热膨胀系数(CTE)是聚合物的一项重要的工业宏观性质。然而,没有基于结构的模型以足够的精度表达它。在这项工作中,我们提出了两个数据驱动的预测模型的线性CTE的无定形均聚物在玻璃态的基础上完全基于化学结构,显示一致的预测。第一个模型是建立与SMILES-X软件,是基于简化的分子输入线输入系统(SMILES)的聚合物的重复单元作为输入。第二个模型是用一个随机森林建立的,该森林是在重复单元的扩展连接指纹上训练的。这两个模型都是在从PoLyInfo数据库中提取的106个实验数据样本上训练的。样本外预测的均方根误差为2.65 ± 0.09 × 10- 5 K-1(2.58 ± 0.09 × 10-5K-1),SMILES-X(随机森林)的平均绝对误差为1.71 ± 0.06 × 10 - 5 K-1(1.61 ± 0.06 × 10- 5 K-1),决定系数为0.62 ± 0.03(0.64 ± 0.03)。此外,该模型进行了实验验证,使用实验室制备的样品具有良好的一致性(p值为两个模型)。SMILES-X中包含的注意力机制指出了SMILES的重要子结构,由此产生的映射表明该模型在化学可解释的基础上做出决策。
The coefficient of thermal expansion (CTE) is an industrially crucial macroscopic property of polymers. Yet, there is no structure-based model expressing it with sufficient accuracy. In this work, we present two data-driven predictive models for the linear CTE of amorphous homopolymers in the glassy state based solely on chemical structure, showing consistent predictions. The first model is built with the SMILES-X software and is based on the simplified molecular-input line-entry system (SMILES) of polymer’s repeating unit as input. The second model is built with a random forest trained on extended-connectivity fingerprints of repeating units. Both models are trained on 106 experimental data samples taken from the PoLyInfo database. The out-of-sample prediction shows a root-mean-square error of 2.65 ± 0.09 × 10–5K–1(2.58 ± 0.09 × 10–5K–1), a mean absolute error of 1.71 ± 0.06 × 10–5K–1(1.61 ± 0.06 × 10–5K–1) and a coefficient of determination of 0.62 ± 0.03 (0.64 ± 0.03) for SMILES-X (random forest). Additionally, the models are validated experimentally using a lab-prepared sample with good agreement (p-valuefor both models). The attention mechanism, incorporated into SMILES-X, points out salient SMILES substructures, and the resulting maps suggest that the model takes decisions on a chemically interpretable basis.Abbreviations:SMILES; CTE; CLTE; CVTE
用于神经架构搜索的无训练模型性能估计
DOI: --
发表时间: 2021
期刊: arXiv.org
影响因子: --
作者:
Ekaterina Gracheva
通讯作者: Ekaterina Gracheva
DOI: 10.1021/ma60036a022
发表时间: 1973
期刊: Macromolecules
影响因子: 5.5
作者:
P. S. Wilson;R. Simha
通讯作者: R. Simha
DOI: 10.1126/sciadv.aaz4301
发表时间: 2020-05-01
期刊: SCIENCE ADVANCES
影响因子: 13.6
作者:
Barnett, J. Wesley;Bilchak, Connor R.;Kumar, Sanat K.
通讯作者: Kumar, Sanat K.
DOI: 10.1021/ma60036a023
发表时间: 1973-01-01
期刊: MACROMOLECULES
影响因子: 5.5
作者:
SIMHA, R;WILSON, PS
通讯作者: WILSON, PS