Cautionary Guidelines for Machine Learning Studies with Combinatorial Datasets

Cautionary Guidelines for Machine Learning Studies with Combinatorial Datasets
复制标题

组合数据集机器学习研究的警示指南

DOI:
10.1021/acscombsci.0c00118
复制
发表时间:
2020
影响因子:
--
通讯作者:
Denmark, Scott E.
Denmark, Scott E.
中科院分区:
化学3区
文献类型:
--
作者:
Zahrt, Andrew F.;Henle, Jeremy J.;Denmark, Scott E.

文献摘要

相似文献

回归建模在有机化学中作为反应结果预测和机理探究的工具变得越来越普遍。通常,为了获得这些研究所需的数据量,研究人员采用组合数据集来最大化数据点的数量,同时限制所需的离散化学实体的数量。在使用组合数据集的建模研究中,一个经常被忽视的问题是倾向于拟合数据集中的模式(即,反应物或催化剂的存在或不存在),而不是识别描述符和响应变量之间的有意义的趋势。因此,这种模型的通用性和可解释性受到影响。本报告在案例研究中说明了这些众所周知的陷阱,演示了必要的控制实验以确定何时此属性会出现问题,并建议如何执行进一步验证以评估使用组合数据集训练的模型的一般适用性和可解释性。
Regression modeling is becoming increasingly prevalent in organic chemistry as a tool for reaction outcome prediction and mechanistic interrogation. Frequently, to acquire the requisite amount of data for such studies, researchers employ combinatorial datasets to maximize the number of data points while limiting the number of discrete chemical entities required. An often-overlooked problem in modeling studies using combinatorial datasets is the tendency to fit on patterns in the datasets (i.e., the presence or absence of a reactant or catalyst) rather than to identify meaningful trends between descriptors and the response variable. Consequently, the generality and interpretability of such models suffer. This report illustrates these well-known pitfalls in a case study, demonstrates the necessary control experiments to identify when this property will be problematic, and suggests how to perform further validation to assess general applicability and interpretability of models trained using combinatorial datasets.