A conceptual study of transfer learning with linear models for data-driven property prediction

A conceptual study of transfer learning with linear models for data-driven property prediction
复制标题

DOI:
10.1016/j.compchemeng.2021.107599
复制
发表时间:
2021-11
期刊:
Comput. Chem. Eng.
影响因子:
--
通讯作者:
Bowen Li;S. Rangarajan
Bowen Li;S. Rangarajan
中科院分区:
其他
文献类型:
--
作者:
Bowen Li;S. Rangarajan

文献摘要

被引文献

相似文献

迁移学习是一个概念,通过共享相关任务的信息,可以为数据可用性有限的任务(例如分子特性)(目标任务)开发数据驱动模型。在化学工程的背景下,这两个任务可以涉及相关属性,也可以涉及以两种不同方式(具有不同的精度或分辨率)计算或测量的相同属性。在这项工作中,我们使用线性和可解释模型的集合提出了一项概念研究,以阐明迁移学习何时有益。我们表明,迁移学习需要两个任务的基础特征有很大的重叠(特别是大于 50%)才能改进目标任务的模型。另一方面,从不相关的任务传输信息(特别是有关显着特征的信息)可能不利于训练目标任务的模型。随后,我们提出了用于分子特性预测的迁移学习的三个说明性示例,并根据我们的概念研究的推论合理化了迁移信息的有用性。因此,这项工作提供了对构建分子属性模型的迁移学习概念的简化分析。
Transfer learning is a concept whereby data-driven models can be developed for tasks (e.g. molecular properties) with limited data availability (target task) by sharing information from a related task. In the context of chemical engineering, the two tasks can either pertain to related properties or to the same property calculated or measured in two different ways (with differing accuracies or resolution). Using an ensemble of linear and interpretable models, in this work, we present a conceptual study to explicate when transfer learning can be beneficial. We show that a large overlap of the underlying features of the two tasks (specifically greater than 50%) is required for transfer learning to improve the model for the target task. On the other hand, transferring information (in particular, information regarding salient features) from an uncorrelated task can be detrimental to train a model for the target task. Subsequently, we present three illustrative examples of transfer learning for molecular property prediction and rationalize the usefulness of transferred information based on the inferences from our conceptual studies. This work, thus, provides a simplified analysis of the concept of transfer learning for building molecular property models.