A conceptual study of transfer learning with linear models for data-driven property prediction
A conceptual study of transfer learning with linear models for data-driven property prediction
复制标题
DOI:
10.1016/j.compchemeng.2021.107599
复制
发表时间:
2021-11
期刊:
影响因子:
--
通讯作者:
Bowen Li;S. Rangarajan
中科院分区:
文献类型:
--
作者:
Bowen Li;S. Rangarajan
Transfer learning is a concept whereby data-driven models can be developed for tasks (e.g. molecular properties) with limited data availability (target task) by sharing information from a related task. In the context of chemical engineering, the two tasks can either pertain to related properties or to the same property calculated or measured in two different ways (with differing accuracies or resolution). Using an ensemble of linear and interpretable models, in this work, we present a conceptual study to explicate when transfer learning can be beneficial. We show that a large overlap of the underlying features of the two tasks (specifically greater than 50%) is required for transfer learning to improve the model for the target task. On the other hand, transferring information (in particular, information regarding salient features) from an uncorrelated task can be detrimental to train a model for the target task. Subsequently, we present three illustrative examples of transfer learning for molecular property prediction and rationalize the usefulness of transferred information based on the inferences from our conceptual studies. This work, thus, provides a simplified analysis of the concept of transfer learning for building molecular property models.