A Precise High-Dimensional Asymptotic Theory for Boosting and Minimum-L1-Norm Interpolated Classifiers

A Precise High-Dimensional Asymptotic Theory for Boosting and Minimum-L1-Norm Interpolated Classifiers
复制标题

Boosting 和最小 L1 范数插值分类器的精确高维渐近理论

DOI:
--
复制
发表时间:
2020
期刊:
Social Science Research Network
影响因子:
--
通讯作者:
P. Sur
P. Sur
中科院分区:
--
文献类型:
--
作者:
Tengyuan Liang;P. Sur

文献摘要

参考文献

被引文献

相似文献

本文从统计和计算的角度出发,建立了一个精确的高维渐近理论,用于可分离数据的提升。我们考虑一个高维设置的功能(弱学习者)的数量$p$规模与样本大小$n$,在过度参数化制度。在一类统计模型下,我们提供了一个精确的分析时,该算法插值的训练数据和最大化的经验$\ell_1$-保证金的推广误差的提升。此外,我们明确地确定了提升测试误差和最佳贝叶斯误差之间的关系,以及插值时活动特征的比例(零初始化)。反过来,这些精确的表征回答了某些问题提出的\cite{breiman 1999 prediction,schapire 1998 boosting}周围的提升,假设的数据生成过程。在我们的理论的核心在于深入研究的最大-$\ell_1$-保证金,这可以准确地描述了一个新的系统的非线性方程,分析这个保证金,我们依靠高斯比较技术,并开发一个新的一致偏差的论点。我们的统计和计算参数可以处理(1)特征分布的任何有限秩尖峰协方差模型和(2)对应于一般$\ell_q$-几何的boosting变体,$q \in [1,2]$。作为最后一个组成部分,通过Lindeberg原则,我们建立了一个普适性结果,显示缩放的$\ell_1 $-保证金(渐近)保持不变,无论用于提升的协变量来自非线性随机特征模型还是具有匹配矩的适当线性化模型。
This paper establishes a precise high-dimensional asymptotic theory for boosting on separable data, taking statistical and computational perspectives. We consider a high-dimensional setting where the number of features (weak learners) $p$ scales with the sample size $n$, in an overparametrized regime. Under a class of statistical models, we provide an exact analysis of the generalization error of boosting when the algorithm interpolates the training data and maximizes the empirical $\ell_1$-margin. Further, we explicitly pin down the relation between the boosting test error and the optimal Bayes error, as well as the proportion of active features at interpolation (with zero initialization). In turn, these precise characterizations answer certain questions raised in \cite{breiman1999prediction, schapire1998boosting} surrounding boosting, under assumed data generating processes. At the heart of our theory lies an in-depth study of the maximum-$\ell_1$-margin, which can be accurately described by a new system of non-linear equations; to analyze this margin, we rely on Gaussian comparison techniques and develop a novel uniform deviation argument. Our statistical and computational arguments can handle (1) any finite-rank spiked covariance model for the feature distribution and (2) variants of boosting corresponding to general $\ell_q$-geometry, $q \in [1, 2]$. As a final component, via the Lindeberg principle, we establish a universality result showcasing that the scaled $\ell_1$-margin (asymptotically) remains the same, whether the covariates used for boosting arise from a non-linear random feature model or an appropriately linearized model with matching moments.
DOI: 10.1080/01621459.2020.1745812
发表时间: 2019-01
影响因子: 3.7
作者:
Xialiang Dou;Tengyuan Liang
通讯作者: Xialiang Dou;Tengyuan Liang
稀疏线性回归斜率的渐近和优化设计
DOI: --
发表时间: 2019
期刊: IEEE International Symposium on Information Theory
影响因子: --
作者:
Hu, Hong;Lu, Yue M.
通讯作者: Lu, Yue M.
DOI: 10.1080/01621459.2016.1273116
发表时间: 2015-10
影响因子: 3.7
作者:
Alexander Hanbo Li;Jelena Bradic
通讯作者: Alexander Hanbo Li;Jelena Bradic
DOI: 10.1007/s00440-018-00896-9
发表时间: 2017-06
影响因子: 2
作者:
P. Sur;Yuxin Chen;E. Candès
通讯作者: P. Sur;Yuxin Chen;E. Candès
DOI: --
发表时间: 2018-06
期刊: ArXiv
影响因子: --
作者:
M. Belkin;A. Rakhlin;A. Tsybakov
通讯作者: M. Belkin;A. Rakhlin;A. Tsybakov