Data-driven ligand field exploration of Fe( iv )–oxo sites for C–H activation

Data-driven ligand field exploration of Fe( iv )–oxo sites for C–H activation
复制标题

用于 C–H 激活的 Fe( iv )–氧代位点的数据驱动配体现场探索

DOI:
10.1039/d2qi01961b
复制
发表时间:
2023
影响因子:
7
通讯作者:
Vogiatzis, Konstantinos D.
Vogiatzis, Konstantinos D.
中科院分区:
化学1区
文献类型:
--
作者:
Jones, Grier M.;Smith, Brett A.;Kirkland, Justin K.;Vogiatzis, Konstantinos D.

文献摘要

相似文献

在酶活性位点中发现的高价 Fe(IV)-oxo 中间体是 C-H 键活化分子催化剂仿生设计的绝佳目标。惰性脂肪烃(例如甲烷)中的 C-H 键具有抵抗化学官能化的强键。为了帮助筛选 C-H 键活化的潜在催化剂,密度泛函理论 (DFT) 和机器学习 (ML) 等计算方法是对广阔的化合物空间进行高通量虚拟搜索的宝贵工具。在这项研究中,我们设计了一个包含 50 种具有不同配位环境的 Fe(IV)-oxo 物种的数据库,这些数据库进一步功能化,总共大约有 181k 个结构。然后对分子数据库的子集进行 DFT 计算,以确定自旋态和 C-H 键激活能。然后根据一系列化学信息标准对收集到的数据进行整理。为了避免对整个化合物空间执行 181k DFT 计算,我们开发了 ML 模型,该模型利用基于持久同源性的新型分子表示,称为持久图像 (PI)。特别是,我们开发了一种新颖的相似性搜索算法,然后训练回归模型来预测 C-H 激活能和分类模型来预测自旋态。首要任务是为 C-H 激活势垒提供高保真度预测。为此,我们将完整数据库分为低保真度结构和高保真度结构,并引入了一个指标 (δΔG‡),该指标评估特定配体修饰相对于母体未取代结构的影响。验证步骤包括对 15 个结构进行额外的 DFT 计算,证明了所提出方法的可信度。
High-valent Fe(IV)–oxo intermediates, found in enzyme active sites, are excellent targets for biomimetic design of molecular catalysts for C–H bond activation. C–H bonds in inert aliphatic hydrocarbons, such as methane, possess strong bonds that are resistant to chemical functionalization. To aid in the screening of potential catalysts for C–H bond activation, computational methods, such as density functional theory (DFT) and machine learning (ML), are valuable tools for performing high-throughput virtual searches of the vast chemical compound space. In this study, we have designed a database of 50 Fe(IV)–oxo species with varying coordination environments which are further functionalized for a total of approximately 181k structures. DFT calculations are then performed on a subset of the molecular database to determine spin states and C–H bond activation energies. The collected data are then curated based on a series of chemically informed criteria. To avoid performing 181k DFT calculations on the total chemical compound space, we developed ML models that utilize a novel molecular representation based on persistence homology, called persistence images (PIs). In particular, we have developed a novel similarity search algorithm, followed by training a regression model to predict C–H activation energies and a classification model to predict the spin states. The priority is to provide high-fidelity predictions for C–H activation barriers. For this purpose, we divided the full database into low- and high-fidelity structures and introduced a metric (δΔG‡) which evaluates the effect of a specific ligand modification with respect to the parent, unsubstituted structure. A validation step that included additional DFT calculations on 15 structures demonstrated the credibility of the proposed methodology.