Reformulating Reactivity Design for Data-Efficient Machine Learning.

Reformulating Reactivity Design for Data-Efficient Machine Learning.
复制标题

DOI:
10.1021/acscatal.3c02513
复制
发表时间:
2023-10-20
期刊:
影响因子:
12.9
通讯作者:
Grayson, Matthew N.
Grayson, Matthew N.
中科院分区:
化学1区
文献类型:
--
作者:
Lewis-Atwell, Toby;Beechey, Daniel;Simsek, Ozgur;Grayson, Matthew N.

文献摘要

参考文献

相似文献

机器学习(ML)可以提供快速、准确的反应障碍预测,用于合理的反应性设计。然而,模型训练需要大数据集,通常是数千或数万个障碍,通过计算或实验获得这些障碍是非常昂贵的。此外,反应空间中每个感兴趣的区域都需要定制的数据集,因为模型通常很难推广。因此,我们将ML障碍预测问题重新表述为一个数据效率更高的过程:从具有所需目标值的预先指定的集合中寻找反应。我们的重新配方能够快速选择具有特定目的激活障碍的反应,例如,在合成的反应性和选择性设计、催化剂设计、毒理学和共价药物发现中,只需要数十个精确测量的障碍。重要的是,我们的重新公式不需要超出手头数据集的范围,并且我们对高度毒理和合成相关的氮杂-迈克尔加成和过渡金属催化的氢气活化的数据集显示了很好的结果,通常需要不到20个精确测量的密度泛函理论(DFT)势垒。即使对于E2和SN2反应的不完整数据集,具有大量缺失障碍(分别为74%和56%),我们选择的ML搜索方法所需的数据点仍然比更传统的ML用于预测激活障碍所需的数百或数千个数据点要少得多。最后,我们包括了一个案例研究,在这个案例中,我们使用我们的过程来指导二氢活化催化剂的优化。我们的方法能够通过只运行12次DFT反应势垒计算来识别目标势垒1千卡摩尔-1内的反应,这说明了这种重新公式对于具有高度合成重要性的体系的使用和现实世界的适用性。
Machine learning (ML) can deliver rapid and accurate reaction barrier predictions for use in rational reactivity design. However, model training requires large data sets of typically thousands or tens of thousands of barriers that are very expensive to obtain computationally or experimentally. Furthermore, bespoke data sets are required for each region of interest in reaction space as models typically struggle to generalize. We have therefore reformulated the ML barrier prediction problem toward a much more data-efficient process: finding a reaction from a prespecified set with a desired target value. Our reformulation enables the rapid selection of reactions with purpose-specific activation barriers, for example, in the design of reactivity and selectivity in synthesis, catalyst design, toxicology, and covalent drug discovery, requiring just tens of accurately measured barriers. Importantly, our reformulation does not require generalization beyond the domain of the data set at hand, and we show excellent results for the highly toxicologically and synthetically relevant data sets of aza-Michael addition and transition-metal-catalyzed dihydrogen activation, typically requiring less than 20 accurately measured density functional theory (DFT) barriers. Even for incomplete data sets of E2 and SN2 reactions, with high numbers of missing barriers (74% and 56% respectively), our chosen ML search method still requires significantly fewer data points than the hundreds or thousands needed for more conventional uses of ML to predict activation barriers. Finally, we include a case study in which we use our process to guide the optimization of the dihydrogen activation catalyst. Our approach was able to identify a reaction within 1 kcal mol–1 of the target barrier by only having to run 12 DFT reaction barrier calculations, which illustrates the usage and real-world applicability of this reformulation for systems of high synthetic importance.
DOI: 10.1039/d0sc00445f
发表时间: 2020-05-14
期刊: Chemical science
影响因子: 8.4
作者:
Friederich P;Dos Passos Gomes G;De Bin R;Aspuru-Guzik A;Balcells D
通讯作者: Balcells D
DOI: 10.1039/d1sc01206a
发表时间: 2021-07-28
期刊: Chemical science
影响因子: 8.4
作者:
Jackson R;Zhang W;Pearson J
通讯作者: Pearson J
DOI: 10.1021/jacs.0c03184
发表时间: 2020-06-24
影响因子: 15
作者:
Bhawal, Benjamin N.;Reisenbauer, Julia C.;Morandi, Bill
通讯作者: Morandi, Bill
DOI: 10.3390/ph15081009
发表时间: 2022-08-17
期刊: Pharmaceuticals (Basel, Switzerland)
影响因子: --
作者:
Cores Á;Clerigué J;Orocio-Rodríguez E;Menéndez JC
通讯作者: Menéndez JC
DOI: 10.1021/acs.jmedchem.8b01153
发表时间: 2019-06-27
影响因子: 7.3
作者:
Gehringer, Matthias;Laufer, Stefan A.
通讯作者: Laufer, Stefan A.