Development of a Computer-Guided Workflow for Catalyst Optimization. Descriptor Validation, Subset Selection, and Training Set Analysis

Development of a Computer-Guided Workflow for Catalyst Optimization. Descriptor Validation, Subset Selection, and Training Set Analysis
复制标题

DOI:
10.1021/jacs.0c04715
复制
发表时间:
2020-07-01
影响因子:
15
通讯作者:
Denmark, Scott E.
Denmark, Scott E.
中科院分区:
化学1区
文献类型:
--
作者:
Henle, Jeremy J.;Zahrt, Andrew F.;Denmark, Scott E.

文献摘要

被引文献

相似文献

现代对映选择性催化剂的发展主要是由经验主义推动的。尽管这种方法促进了大多数现有合成方法的引入,但它本质上受到实践者的技能、创造力和化学直觉的限制。在这里,我们提出了一种互补的方法来催化剂优化,其中统计方法在每个阶段使用,以简化发展。为了构建优化信息学工作流,必须对许多关键组件进行严格的验证。首先,在两个案例研究中验证了至关重要的分子描述符,以确定构象依赖分子表征的重要性。接下来,有了大量可用的数据集,就有可能研究使用不同建模方法建立预测模型所需的数据量。考虑到许多催化剂结构的商业可用性,可以将算法选择的训练集和商业可用的训练集生成的模型进行比较。最后,通过无监督学习的方法证明了有限数据集的增强,以恢复生成模型的准确性。
Modern, enantioselective catalyst development is driven largely by empiricism. Although this approach has fostered the introduction of most of the existing synthetic methods, it is inherently limited by the skill, creativity, and chemical intuition of the practitioner. Herein, we present a complementary approach to catalyst optimization in which statistical methods are used at each stage to streamline development. To construct the optimization informatics workflow, a number of critical components had to be subjected to rigorous validation. First, the critically important molecular descriptors were validated in two case studies to establish the importance of conformation-dependent molecular representations. Next, with a large data set available, it was possible to investigate the amount of data necessary to make predictive models with different modeling methods. Given the commercial availability of many catalyst structures, it was possible to compare models generated with algorithmically selected training sets and commercially available training sets. Finally, the augmentation of limited data sets is demonstrated in a method informed by unsupervised learning to restore the accuracy of the generated models.