Gradient descent for robust kernel-based regression

Gradient descent for robust kernel-based regression
复制标题

用于稳健的基于核的回归的梯度下降

DOI:
10.1088/1361-6420/aabe55
复制
发表时间:
2018-06-01
期刊:
影响因子:
2.1
通讯作者:
Shi, Lei
Shi, Lei
中科院分区:
数学2区
文献类型:
--
作者:
Guo, Zheng-Chu;Hu, Ting;Shi, Lei

文献摘要

被引文献

相似文献

本文研究了再生核Hilbert空间(RKHS)上由鲁棒损失函数l(sigma)生成的梯度下降算法。损失函数由窗口函数G和尺度参数sigma定义,其可以包括广泛的常用回归鲁棒损失。基于l(sigma)损失的经验风险最小化问题的理论分析与优化过程之间仍存在差距:理论分析中需要估计量全局最优,而优化方法不能保证其解的全局最优性。在本文中,我们的目标是填补这一空白,通过开发一种新的理论分析的性能产生的梯度下降算法的估计。我们证明,适当选择尺度参数σ,梯度更新与早期停止规则可以近似回归函数。我们优雅的误差分析可以导致标准L-2范数和强RKHS范数的收敛,这两者在最小-最大意义下都是最优的。我们发现,尺度参数σ在提供鲁棒性和快速收敛方面起着重要的作用。在合成的例子和真实的数据集上进行的数值实验也支持了我们的理论结果。
In this paper, we study the gradient descent algorithm generated by a robust loss function l(sigma) over a reproducing kernel Hilbert space (RKHS). The loss function is defined by a windowing function G and a scale parameter sigma, which can include a wide range of commonly used robust losses for regression. There is still a gap between theoretical analysis and optimization process of empirical risk minimization based on l(sigma) loss: the estimator needs to be global optimal in the theoretical analysis while the optimization method can not ensure the global optimality of its solutions. In this paper, we aim to fill this gap by developing a novel theoretical analysis on the performance of estimators generated by the gradient descent algorithm. We demonstrate that with an appropriately chosen scale parameter sigma, the gradient update with early stopping rules can approximate the regression function. Our elegant error analysis can lead to convergence in the standard L-2 norm and the strong RKHS norm, both of which are optimal in the mini-max sense. We show that the scale parameter sigma plays an important role in providing robustness as well as fast convergence. The numerical experiments implemented on synthetic examples and real data set also support our theoretical results.