Featurization strategies for polymer sequence or composition design by machine learning

Featurization strategies for polymer sequence or composition design by machine learning
复制标题

通过机器学习进行聚合​​物序列或组合物设计的特征化策略

DOI:
10.1039/d1me00160d
复制
发表时间:
2022
影响因子:
3.6
通讯作者:
Webb, Michael A.
Webb, Michael A.
中科院分区:
工程技术3区
文献类型:
--
作者:
Patel, Roshan A.;Borca, Carlos H.;Webb, Michael A.

文献摘要

相似文献

数据密集型科学发现和机器学习的出现极大地改变了科学家和工程师处理材料设计的方式。然而,对于设计大分子或聚合物,一个限制是缺乏适当的方法或标准,用于将系统转换为化学信息,机器可读的表示。这种特征化过程对于构建可以指导聚合物发现的预测模型至关重要。虽然标准的分子特征化技术已经部署在均聚物上,但这种方法既不能捕获共聚物的多尺度性质,也不能捕获共聚物的拓扑复杂性,并且它们对不能由单个重复单元表征的系统的应用有限。在这里,我们提出,评估和分析了一系列的特征化策略,适用于共聚物系统。这些策略在来自四个不同数据集的不同预测任务中进行了系统的检查,这些数据集使我们能够了解特征化如何影响共聚物性能预测。基于这种比较分析,我们建议在可能的情况下直接在聚合物表示中编码聚合物尺寸,在精确的聚合物序列已知时采用拓扑描述符或卷积神经网络,并在开发外推模型时使用化学信息单元表示。这些结果为通过机器学习进行共聚物设计的聚合物特征化提供了指导和未来的方向。
The emergence of data-intensive scientific discovery and machine learning has dramatically changed the way in which scientists and engineers approach materials design. Nevertheless, for designing macromolecules or polymers, one limitation is the lack of appropriate methods or standards for converting systems into chemically informed, machine-readable representations. This featurization process is critical to building predictive models that can guide polymer discovery. Although standard molecular featurization techniques have been deployed on homopolymers, such approaches capture neither the multiscale nature nor topological complexity of copolymers, and they have limited application to systems that cannot be characterized by a single repeat unit. Herein, we present, evaluate, and analyze a series of featurization strategies suitable for copolymer systems. These strategies are systematically examined in diverse prediction tasks sourced from four distinct datasets that enable understanding of how featurization can impact copolymer property prediction. Based on this comparative analysis, we suggest directly encoding polymer size in polymer representations when possible, adopting topological descriptors or convolutional neural networks when the precise polymer sequence is known, and using chemically informed unit representations when developing extrapolative models. These results provide guidance and future directions regarding polymer featurization for copolymer design by machine learning.