Data Augmentation for Support Vector Machines

Data Augmentation for Support Vector Machines
复制标题

DOI:
10.1214/11-ba601
复制
发表时间:
2011-01-01
期刊:
影响因子:
4.4
通讯作者:
Scott, Steven L.
Scott, Steven L.
中科院分区:
数学2区
文献类型:
--
作者:
Polson, Nicholas G.;Scott, Steven L.

文献摘要

被引文献

相似文献

本文提出了正则化支持向量机 (SVM) 的潜在变量表示,使 EM、ECME 或 MCMC 算法能够提供参数估计。我们通过证明最小化 SVM 最优性标准和参数正则化惩罚等效于找到正态伪后验分布的均值方差混合模式来验证我们的表示。混合表示中的潜在变量导致 SVM 参数的 EM 和 ECME 点估计,以及基于吉布斯采样的 MCMC 算法,该算法可以将高斯线性模型的贝叶斯工具应用于 SVM。我们展示了如何使用尖峰和平板先验实现 SVM,并针对来自标准垃圾邮件过滤数据集的数据运行它们。
This paper presents a latent variable representation of regularized support vector machines (SVM's) that enables EM, ECME or MCMC algorithms to provide parameter estimates. We verify our representation by demonstrating that minimizing the SVM optimality criterion together with the parameter regularization penalty is equivalent to finding the mode of a mean-variance mixture of normals pseudo-posterior distribution. The latent variables in the mixture representation lead to EM and ECME point estimates of SVM parameters, as well as MCMC algorithms based on Gibbs sampling that can bring Bayesian tools for Gaussian linear models to bear on SVM's. We show how to implement SVM's with spike-and-slab priors and run them against data from a standard spam filtering data set.