Excess risk bounds in robust empirical risk minimization

Excess risk bounds in robust empirical risk minimization
复制标题

DOI:
10.1093/imaiai/iaab004
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Stanislav Minsker;Timothée Mathieu
Stanislav Minsker;Timothée Mathieu
中科院分区:
其他
文献类型:
--
作者:
Stanislav Minsker;Timothée Mathieu

文献摘要

被引文献

相似文献

本文研究了通用经验风险最小化算法的稳健版本,这是现代统计方法的核心技术之一。经验风险最小化的成功基于以下事实:对于“行为良好”的随机过程 $\left \{ f(X), \ f\in \mathscr F\right \}$ 由一类函数 $f\in \mathscr F$ 索引,在样本 $X_1,\ldots 上评估平均值 $\frac{1}{N}\sum _{j=1}^N f(X_j)$ ,独立同分布的 X_N$ $X$ 的副本提供了对期望 $\mathbb E f(X)$ 的良好近似,在大类 $f\in \mathscr F$ 上一致。然而,如果过程的边际分布是重尾的或者样本包含异常值,则情况可能不再成立。我们提出了一种经验风险最小化的版本,其基础是用稳健的期望代理替换样本平均值,并获得结果估计量的超额风险的高置信界限。特别是,我们表明,相对于样本大小 $N$,鲁棒估计量的超额风险可以快速收敛到 $0$,指的是比 $N^{-1/2}$ 更快的速率。我们讨论主要结果对线性和逻辑回归问题的影响,并评估所提出的方法在模拟和实际数据上的数值性能。
This paper investigates robust versions of the general empirical risk minimization algorithm, one of the core techniques underlying modern statistical methods. Success of the empirical risk minimization is based on the fact that for a ‘well-behaved’ stochastic process $\left \{ f(X), \ f\in \mathscr F\right \}$ indexed by a class of functions $f\in \mathscr F$, averages $\frac{1}{N}\sum _{j=1}^N f(X_j)$ evaluated over a sample $X_1,\ldots ,X_N$ of i.i.d. copies of $X$ provide good approximation to the expectations $\mathbb E f(X)$, uniformly over large classes $f\in \mathscr F$. However, this might no longer be true if the marginal distributions of the process are heavy tailed or if the sample contains outliers. We propose a version of empirical risk minimization based on the idea of replacing sample averages by robust proxies of the expectations and obtain high-confidence bounds for the excess risk of resulting estimators. In particular, we show that the excess risk of robust estimators can converge to $0$ at fast rates with respect to the sample size $N$, referring to the rates faster than $N^{-1/2}$. We discuss implications of the main results to the linear and logistic regression problems and evaluate the numerical performance of proposed methods on simulated and real data.