Supervised learning via Euler's Elastica models

Supervised learning via Euler's Elastica models
复制标题

DOI:
10.5555/2789272.2912113
复制
发表时间:
2015
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Tong Lin;Hanlin Xue;Ling Wang;Bo Huang;H. Zha
Tong Lin;Hanlin Xue;Ling Wang;Bo Huang;H. Zha
中科院分区:
其他
文献类型:
--
作者:
Tong Lin;Hanlin Xue;Ling Wang;Bo Huang;H. Zha

文献摘要

相似文献

本文研究了函数逼近框架下高维监督学习问题的欧拉弹性(EE)模型。1744年,欧拉在无扭弹性细杆的建模中引入了二维曲线的弹性能量。在过去的二十年里,欧拉弹性函数与其退化形式的全变分(TV)一起被成功地应用于低维数据处理,如图像去噪和图像修复。我们的动机是应用欧拉的弹性高维监督学习问题。为此,监督学习问题被建模为一个新的几何正则化方案下的能量泛函最小化,其中的能量是由平方损失和弹性惩罚。elastica罚分旨在通过对所有水平曲线上的大梯度和高曲率值进行严重惩罚来正则化近似函数。我们采用计算偏微分方程的方法来最小化能量泛函。利用变分原理,将能量最小化问题转化为欧拉-拉格朗日偏微分方程。然而,这种偏微分方程通常是高维的,无法由常见的低维求解器直接处理。为了克服这个困难,我们使用径向基函数(RBF)来近似目标函数,这减少了优化问题,找到这些基函数的线性系数。分析了该模型解的存在唯一性和泛相容性等理论性质。大量的实验已经证明了该模型的有效性,二进制分类,多类分类,和回归任务。
This paper investigates the Euler's elastica (EE) model for high-dimensional supervised learning problems in a function approximation framework. In 1744 Euler introduced the elastica energy for a 2D curve on modeling torsion-free thin elastic rods. Together with its degenerate form of total variation (TV), Euler's elastica has been successfully applied to low-dimensional data processing such as image denoising and image inpainting in the last two decades. Our motivation is to apply Euler's elastica to high-dimensional supervised learning problems. To this end, a supervised learning problem is modeled as an energy functional minimization under a new geometric regularization scheme, where the energy is composed of a squared loss and an elastica penalty. The elastica penalty aims at regularizing the approximated function by heavily penalizing large gradients and high curvature values on all level curves. We take a computational PDE approach to minimize the energy functional. By using variational principles, the energy minimization problem is transformed into an Euler-Lagrange PDE. However, this PDE is usually high-dimensional and can not be directly handled by common low-dimensional solvers. To circumvent this difficulty, we use radial basis functions (RBF) to approximate the target function, which reduces the optimization problem to finding the linear coefficients of these basis functions. Some theoretical properties of this new model, including the existence and uniqueness of solutions and universal consistency, are analyzed. Extensive experiments have demonstrated the effectiveness of the proposed model for binary classification, multi-class classification, and regression tasks.