Aggregation for Gaussian regression

Aggregation for Gaussian regression
复制标题

DOI:
10.1214/009053606000001587
复制
发表时间:
2007-08
影响因子:
4.5
通讯作者:
F. Bunea;A. Tsybakov;M. Wegkamp
F. Bunea;A. Tsybakov;M. Wegkamp
中科院分区:
数学1区
文献类型:
--
作者:
F. Bunea;A. Tsybakov;M. Wegkamp

文献摘要

被引文献

相似文献

本文研究了回归环境中的统计聚集过程。一个激励因素是存在许多不同的估计方法,导致可能相互竞争的估计。我们在这里考虑三种不同类型的聚集:模型选择(MS)聚集,凸(C)聚集和线性(L)聚集。(MS)的目标是从列表中选择最优的单个估计量;(C)的目标是选择给定估计量的最优凸组合;(L)的目标是选择给定估计量的最优线性组合。我们感兴趣的是评估这些程序获得的估计的超额风险的收敛速度。我们的方法是由最近发表的极大极小结果[Nemirovski,A。(2000年)的第10/2000号决议。非参数统计中的主题。概率论与统计讲座(圣面粉,1998年)。数学讲义。1738年85-277。Springer,柏林; Tsybakov,A. B。(2003年)的报告。最佳聚合率。学习理论与核心机器人工智能讲义2777 303-313. Springer,Heidelberg].存在分别针对(MS)、(C)和(L)情况中的每一个实现最优收敛速率的竞争聚集过程。由于这些程序不能直接相互比较,我们建议一个替代的解决方案。我们证明,所有三个最佳速率,以及那些新引入的(S)聚合(子集选择),几乎是通过一个单一的“通用”的聚合过程。该过程包括混合的初始估计与惩罚最小二乘获得的权重。两种不同的惩罚被认为是:其中之一是BIC型,第二个是一个数据相关的l1型惩罚。
This paper studies statistical aggregation procedures in the regression setting. A motivating factor is the existence of many different methods of estimation, leading to possibly competing estimators. We consider here three different types of aggregation: model selection (MS) aggregation, convex (C) aggregation and linear (L) aggregation. The objective of (MS) is to select the optimal single estimator from the list; that of (C) is to select the optimal convex combination of the given estimators; and that of (L) is to select the optimal linear combination of the given estimators. We are interested in evaluating the rates of convergence of the excess risks of the estimators obtained by these procedures. Our approach is motivated by recently published minimax results [Nemirovski, A. (2000). Topics in non-parametric statistics. Lectures on Probability Theory and Statistics (Saint-Flour, 1998). Lecture Notes in Math. 1738 85-277. Springer, Berlin; Tsybakov, A. B. (2003). Optimal rates of aggregation. Learning Theory and Kernel Machines. Lecture Notes in Artificial Intelligence 2777 303-313. Springer, Heidelberg]. There exist competing aggregation procedures achieving optimal convergence rates for each of the (MS), (C) and (L) cases separately. Since these procedures are not directly comparable with each other, we suggest an alternative solution. We prove that all three optimal rates, as well as those for the newly introduced (S) aggregation (subset selection), are nearly achieved via a single "universal" aggregation procedure. The procedure consists of mixing the initial estimators with weights obtained by penalized least squares. Two different penalties are considered: one of them is of the BIC type, the second one is a data-dependent l 1 -type penalty.