How Data Augmentation affects Optimization for Linear Regression

How Data Augmentation affects Optimization for Linear Regression
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
B. Hanin;Yi Sun
B. Hanin;Yi Sun
中科院分区:
其他
文献类型:
--
作者:
B. Hanin;Yi Sun

文献摘要

被引文献

相似文献

虽然数据增强已迅速成为现代机器学习中优化的关键工具,但增强计划如何影响优化并与优化超参数(如学习率)相互作用的清晰画面刚刚出现。在经典凸优化和最近的工作隐式偏差的精神,目前的工作分析了增广优化的效果在简单的凸设置的线性回归MSE损失。我们找到了学习率和数据增强方案的联合时间表,在该方案下,增强梯度下降可证明收敛,并描述了所得的最小值。我们的研究结果适用于任意增强方案,揭示了复杂的学习率和增强之间的相互作用,即使在凸设置。我们的方法解释增广(S)GD作为一个随机优化方法的代理损失随时间变化的序列。这提供了一种统一的方法来分析学习率,批量大小和从加性噪声到随机投影的增强。从这个角度来看,我们的结果,这也给出了收敛速度,可以被视为Monro-Robbins型条件增广(S)GD。
Though data augmentation has rapidly emerged as a key tool for optimization in modern machine learning, a clear picture of how augmentation schedules affect optimization and interact with optimization hyperparameters such as learning rate is nascent. In the spirit of classical convex optimization and recent work on implicit bias, the present work analyzes the effect of augmentation on optimization in the simple convex setting of linear regression with MSE loss. We find joint schedules for learning rate and data augmentation scheme under which augmented gradient descent provably converges and characterize the resulting minimum. Our results apply to arbitrary augmentation schemes, revealing complex interactions between learning rates and augmentations even in the convex setting. Our approach interprets augmented (S)GD as a stochastic optimization method for a time-varying sequence of proxy losses. This gives a unified way to analyze learning rate, batch size, and augmentations ranging from additive noise to random projections. From this perspective, our results, which also give rates of convergence, can be viewed as Monro-Robbins type conditions for augmented (S)GD.