Benchmarking Machine Learning Models for Polymer Informatics: An Example of Glass Transition Temperature

Benchmarking Machine Learning Models for Polymer Informatics: An Example of Glass Transition Temperature
复制标题

DOI:
10.1021/acs.jcim.1c01031
复制
发表时间:
2021-11-22
影响因子:
5.6
通讯作者:
Li, Ying
Li, Ying
中科院分区:
化学2区
文献类型:
--
作者:
Tao, Lei;Varshney, Vikas;Li, Ying

文献摘要

被引文献

相似文献

在聚合物信息学领域,利用机器学习(ML)技术来评估聚合物的玻璃化转变温度T-g等特性引起了广泛的关注。当遇到数量惊人的聚合物结构时,这种以数据为中心的方法比费力的实验测量更有效和实用。各种ML模型被证明在T-g预测方面表现良好。然而,它们是在不同的数据集上训练的,使用不同的结构表示,并基于不同的特征工程方法。因此,选择合适的ML模型来更好地处理具有泛化能力的T-g预测就成为关键问题。为了对不同的机器学习技术进行公平的比较,并检查影响模型性能的关键因素,我们通过编译79种不同的机器学习模型并在大型和多样化的数据集上训练它们,进行了系统的基准研究。建立ML模型的三个主要组成部分是结构表示、特征表示和ML算法。在聚合物结构表征方面,我们考虑了具有较长链结构的聚合物单体、重复单元和低聚物。基于该特征,计算表征,包括有或没有子结构频率的摩根指纹、RDKit描述符、分子嵌入、分子图等。然后,使用不同的机器学习算法,如深度神经网络、卷积神经网络、随机森林、支持向量机、LASSO回归和高斯过程回归,对得到的特征输入进行训练。我们使用holdout测试集和来自高通量分子动力学模拟的额外未标记数据集来评估这些ML模型的性能。特别关注了机器学习模型在未标记数据集上的泛化能力,并考虑了模型对拓扑结构和聚合物分子量的敏感性。该基准研究不仅为Tg预测任务提供了指导,也为其他聚合物信息学任务提供了有益的参考。
In the field of polymer informatics, utilizing machine learning (ML) techniques to evaluate the glass transition temperature T-g and other properties of polymers has attracted extensive attention. This data-centric approach is much more efficient and practical than the laborious experimental measurements when encountered a daunting number of polymer structures. Various ML models are demonstrated to perform well for T-g prediction. Nevertheless, they are trained on different data sets, using different structure representations, and based on different feature engineering methods. Thus, the critical question arises on selecting a proper ML model to better handle the T-g prediction with generalization ability. To provide a fair comparison of different ML techniques and examine the key factors that affect the model performance, we carry out a systematic benchmark study by compiling 79 different ML models and training them on a large and diverse data set. The three major components in setting up an ML model are structure representations, feature representations, and ML algorithms. In terms of polymer structure representation, we consider the polymer monomer, repeat unit, and oligomer with longer chain structure. Based on that feature, representation is calculated, including Morgan fingerprinting with or without substructure frequency, RDKit descriptors, molecular embedding, molecular graph, etc. Afterward, the obtained feature input is trained using different ML algorithms, such as deep neural networks, convolutional neural networks, random forest, support vector machine, LASSO regression, and Gaussian process regression. We evaluate the performance of these ML models using a holdout test set and an extra unlabeled data set from high-throughput molecular dynamics simulation. The ML model's generalization ability on an unlabeled data set is especially focused, and the model's sensitivity to topology and the molecular weight of polymers is also taken into consideration. This benchmark study provides not only a guideline for the Tg prediction task but also a useful reference for other polymer informatics tasks.