Learning Augmentation Distributions using Transformed Risk Minimization

Learning Augmentation Distributions using Transformed Risk Minimization
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Evangelos Chatzipantazis;Stefanos Pertigkiozoglou;Edgar Dobriban;Kostas Daniilidis
Evangelos Chatzipantazis;Stefanos Pertigkiozoglou;Edgar Dobriban;Kostas Daniilidis
中科院分区:
其他
文献类型:
--
作者:
Evangelos Chatzipantazis;Stefanos Pertigkiozoglou;Edgar Dobriban;Kostas Daniilidis

文献摘要

相似文献

我们提出了一个新的 \emph{Transformed Risk Minimization} (TRM) 框架作为经典风险最小化的扩展。在TRM中,我们不仅优化预测模型,还优化数据转换;特别是在其分布上。作为一个关键应用,我们专注于学习增强;例如,适当旋转图像,以提高给定类别预测变量的分类性能。我们的 TRM 方法 (1) 在 \emph{单训练循环} 中联合学习变换和模型,(2) 与适用于标准风险最小化的任何训练算法一起使用,(3) 处理任何变换,例如离散和连续类别的增强。为了避免在实施经验转换风险最小化时过度拟合,我们提出了一种基于 PAC-Bayes 理论的新型正则化器。为了学习图像增强,我们提出了通过几何变换块的随机组合来对增强空间进行新的参数化。这导致了新的 \emph{随机组合增强学习} (SCALE) 算法。使用 SCALE 的 TRM 的性能优于 CIFAR10/100 上的先前方法。此外,我们凭经验证明 SCALE 可以正确学习数据分布中的某些对称性(在旋转的 MNIST 上恢复旋转),并且还可以改进学习模型的校准。
We propose a new \emph{Transformed Risk Minimization} (TRM) framework as an extension of classical risk minimization. In TRM, we optimize not only over predictive models, but also over data transformations; specifically over distributions thereof. As a key application, we focus on learning augmentations; for instance appropriate rotations of images, to improve classification performance with a given class of predictors. Our TRM method (1) jointly learns transformations and models in a \emph{single training loop}, (2) works with any training algorithm applicable to standard risk minimization, and (3) handles any transforms, such as discrete and continuous classes of augmentations. To avoid overfitting when implementing empirical transformed risk minimization, we propose a novel regularizer based on PAC-Bayes theory. For learning augmentations of images, we propose a new parametrization of the space of augmentations via a stochastic composition of blocks of geometric transforms. This leads to the new \emph{Stochastic Compositional Augmentation Learning} (SCALE) algorithm. The performance of TRM with SCALE compares favorably to prior methods on CIFAR10/100. Additionally, we show empirically that SCALE can correctly learn certain symmetries in the data distribution (recovering rotations on rotated MNIST) and can also improve calibration of the learned model.