Classification vs regression in overparameterized regimes: Does the loss function matter?

Classification vs regression in overparameterized regimes: Does the loss function matter?
复制标题

DOI:
--
复制
发表时间:
2020-05
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Vidya Muthukumar;Adhyyan Narang;Vignesh Subramanian;M. Belkin;Daniel J. Hsu;A. Sahai
Vidya Muthukumar;Adhyyan Narang;Vignesh Subramanian;M. Belkin;Daniel J. Hsu;A. Sahai
中科院分区:
其他
文献类型:
--
作者:
Vidya Muthukumar;Adhyyan Narang;Vignesh Subramanian;M. Belkin;Daniel J. Hsu;A. Sahai

文献摘要

被引文献

相似文献

我们将过度参数化线性模型中的分类和回归任务与高斯特征进行比较。一方面,我们表明,有足够的过度参数化,所有训练点都是支持向量:最小二乘最小值最小值插值获得的解决方案,通常用于回归,与Hard-Margin支持矢量机(SVM)产生的解决方案相同。这可以最大程度地减少铰链损失,通常用于训练分类器。另一方面,我们表明存在通过0-1测试损耗函数评估这些解决方案几乎最佳的制度,但是如果通过正方形损耗函数进行评估,则不会概括,即它们实现无效的风险。我们的结果表明,在训练阶段(优化)和测试阶段(概括)中使用的损失函数的作用和特性非常不同。
We compare classification and regression tasks in the overparameterized linear model with Gaussian features. On the one hand, we show that with sufficient overparameterization all training points are support vectors: solutions obtained by least-squares minimum-norm interpolation, typically used for regression, are identical to those produced by the hard-margin support vector machine (SVM) that minimizes the hinge loss, typically used for training classifiers. On the other hand, we show that there exist regimes where these solutions are near-optimal when evaluated by the 0-1 test loss function, but do not generalize if evaluated by the square loss function, i.e. they achieve the null risk. Our results demonstrate the very different roles and properties of loss functions used at the training phase (optimization) and the testing phase (generalization).