Interpretable Machine Learning for Catalytic Materials Design toward Sustainability

Interpretable Machine Learning for Catalytic Materials Design toward Sustainability
复制标题

DOI:
10.1021/accountsmr.3c00131
复制
发表时间:
2023-11
影响因子:
14.6
通讯作者:
Hongliang Xin;Tianyou Mou;H. Pillai;Shih-Han Wang;Yang Huang
Hongliang Xin;Tianyou Mou;H. Pillai;Shih-Han Wang;Yang Huang
中科院分区:
--
文献类型:
--
作者:
Hongliang Xin;Tianyou Mou;H. Pillai;Shih-Han Wang;Yang Huang

文献摘要

相似文献

概述寻找具有最佳性能的催化材料以实现可持续化学和能源转化是当今社会面临的紧迫挑战之一。传统上,催化剂或炼金术士的点金石的发现依赖于物理化学直觉的试错方法。科学和工程领域,特别是量子化学和计算基础设施领域数十年的进步,普及了材料发现的计算科学范式。然而,在广阔的化学空间中进行强力搜索因其巨大的成本而受到阻碍。近年来,机器学习 (ML) 已成为一种通过从数据中学习来简化活动站点设计的有前途的方法。随着机器学习越来越多地用于在实际环境中进行预测,对领域可解释性的需求正在激增。因此,深入回顾我们在解决计算多相催化这一挑战性问题方面所做的努力非常重要。在本篇文章中,我们提出了一个可解释的机器学习框架,用于加速催化材料设计,特别是在驱动可持续的碳、氮和氧循环方面。通过利用线性吸附-能量标度和布朗斯台德-埃文斯-波兰尼 (BEP) 关系,多步反应的催化结果(即活性、选择性和稳定性)通常可以映射到一两个动力学相关描述符上。一种非常重要的描述符是活性位点基序上代表性物种的吸附能,可以通过量子化学模拟计算。为了补充这种基于描述符的设计策略,我们描述了将领域知识纳入数据驱动的机器学习工作流程的努力。我们证明,黑盒机器学习算法的主要缺点(例如可解释性差)可以通过采用(1)物理启发的特征工程、(2)贝叶斯统计学习和(3)注入理论的深度神经网络来很大程度上规避。该框架极大地促进了非均相金属基催化剂的设计,其中一些催化剂已经过一系列可持续化学物质的实验验证。我们对可解释机器学习在预测催化材料方面的现有挑战、机遇和未来方向进行了一些评论,更重要的是,对超越传统智慧的催化理论进行了一些评论。我们预计该账户将吸引更多研究人员的注意力,以开发高度准确、易于解释且值得信赖的材料设计策略,通过催化促进向可持续发展的数据科学范式的过渡。
ConspectusFinding catalytic materials with optimal properties for sustainable chemical and energy transformations is one of the pressing challenges facing our society today. Traditionally, the discovery of catalysts or the philosopher’s stone of alchemists relies on a trial-and-error approach with physicochemical intuition. Decades-long advances in science and engineering, particularly in quantum chemistry and computing infrastructures, popularize a paradigm of computational science for materials discovery. However, the brute-force search through a vast chemical space is hampered by its formidable cost. In recent years, machine learning (ML) has emerged as a promising approach to streamline the design of active sites by learning from data. As ML is increasingly employed to make predictions in practical settings, the demand for domain interpretability is surging. Therefore, it is of great importance to provide an in-depth review of our efforts in tackling this challenging issue in computational heterogeneous catalysis.In this Account, we present an interpretable ML framework for accelerating catalytic materials design, particularly in driving sustainable carbon, nitrogen, and oxygen cycles. By leveraging the linear adsorption-energy scaling and Brønsted–Evans–Polanyi (BEP) relationships, catalytic outcomes (i.e., activity, selectivity, and stability) of a multistep reaction can often be mapped onto one or two kinetics-informed descriptors. One type of descriptor of great importance is the adsorption energies of representative species at active site motifs that can be computed from quantum-chemical simulations. To complement such a descriptor-based design strategy, we delineate our endeavors in incorporating domain knowledge into a data-driven ML workflow. We demonstrate that the major drawbacks of black-box ML algorithms, e.g., poor explainability, can be largely circumvented by employing (1) physics-inspired feature engineering, (2) Bayesian statistical learning, and (3) theory-infused deep neural networks. The framework drastically facilitates the design of heterogeneous metal-based catalysts, some of which have been experimentally verified for an array of sustainable chemistries. We offer some remarks on the existing challenges, opportunities, and future directions of interpretable ML in predicting catalytic materials and, more importantly, on advancing catalysis theory beyond conventional wisdom. We envision that this Account will attract more researchers’ attention to develop highly accurate, easily explainable, and trustworthy materials design strategies, facilitating the transition to the data science paradigm for sustainability through catalysis.