Support vector machines for predictive modeling in heterogeneous catalysis: A comprehensive introduction and overfitting investigation based on two real applications

Support vector machines for predictive modeling in heterogeneous catalysis: A comprehensive introduction and overfitting investigation based on two real applications
复制标题

DOI:
10.1021/cc050093m
复制
发表时间:
2006-07-10
影响因子:
--
通讯作者:
Corma, A.
Corma, A.
中科院分区:
其他
文献类型:
--
作者:
Baumes, L. A.;Serra, J. M.;Corma, A.

文献摘要

被引文献

相似文献

本文介绍了用于多相催化预测建模的支持向量机(SVMs),逐步描述了该方法,并重点介绍了该技术成为一种有吸引力的方法的要点。我们首先研究线性支持向量机,通过一个简单的例子详细地工作,该例子基于一项旨在应用高通量实验优化烯烃环氧化催化剂的实验数据。由于所研究的催化变量很少,因此选择了这个案例来直观地强调支持向量机的特征。展示了支持向量机如何将原始数据转换到另一个更高维的表示空间。引入了Vapnik-Chervonenkis维数和结构风险最小化的概念。对支持向量机方法进行了第二个催化应用,即轻烃异构化。最后,我们讨论了为什么与其他机器学习技术,如神经网络或归纳树相比,支持向量机是一种战略方法,以及为什么强调过拟合问题。
This works provides an introduction to support vector machines (SVMs) for predictive modeling in heterogeneous catalysis, describing step by step the methodology with a highlighting of the points which make such technique an attractive approach. We first investigate linear SVMs, working in detail through a simple example based on experimental data derived from a study aiming at optimizing olefin epoxidation catalysts applying high-throughput experimentation. This case study has been chosen to underline SVM features in a visual manner because of the few catalytic variables investigated. It is shown how SVMs transform original data into another representation space of higher dimensionality. The concepts of Vapnik-Chervonenkis dimension and structural risk minimization are introduced. The SVM methodology is evaluated with a second catalytic application, that is, light paraffin isomerization. Finally, we discuss why SVMs is a strategic method, as compared to other machine learning techniques, such as neural networks or induction trees, and why emphasis is put on the problem of overfitting.