A Tunable Loss Function for Robust Classification: Calibration, Landscape, and Generalization

A Tunable Loss Function for Robust Classification: Calibration, Landscape, and Generalization
复制标题

DOI:
10.1109/tit.2022.3169440
复制
发表时间:
2019-06
影响因子:
2.5
通讯作者:
Tyler Sypherd;Mario Díaz;J. Cava;Gautam Dasarathy;P. Kairouz;L. Sankar
Tyler Sypherd;Mario Díaz;J. Cava;Gautam Dasarathy;P. Kairouz;L. Sankar
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tyler Sypherd;Mario Díaz;J. Cava;Gautam Dasarathy;P. Kairouz;L. Sankar

文献摘要

被引文献

相似文献

我们引入了一个可调的损失函数$\alpha $ -loss,参数化为$\alpha \in (0,\infty]$),它在指数损失($\alpha = 1/2$)、对数损失($\alpha = 1$)和0 - 1损失($\alpha = \infty $)之间进行插值,用于机器学习的分类设置。从理论上讲,我们说明了$\alpha $ -loss和Arimoto条件熵之间的基本联系,验证了$\alpha $ -loss的分类校准,以便通过Rademacher复杂度泛化技术证明渐近最优性,并建立了一个称为严格局部拟凸的概念,以便定量表征$\alpha $ -loss的优化景观。实际上,我们使用卷积神经网络对基准图像数据集进行了类不平衡、鲁棒性和分类实验。我们的主要实际结论是,将$\alpha $ -loss从log-loss ($\alpha = 1$)中调优可能会使某些任务受益,为此我们为实践者提供了简单的启发式方法。特别是,导航$\alpha $超参数可以很容易地为标签翻转($\alpha > 1$)提供卓越的模型鲁棒性和对不平衡类($\alpha)的灵敏度。
We introduce a tunable loss function called $\alpha $ -loss, parameterized by $\alpha \in (0,\infty]$ , which interpolates between the exponential loss ( $\alpha = 1/2$ ), the log-loss ( $\alpha = 1$ ), and the 0–1 loss ( $\alpha = \infty $ ), for the machine learning setting of classification. Theoretically, we illustrate a fundamental connection between $\alpha $ -loss and Arimoto conditional entropy, verify the classification-calibration of $\alpha $ -loss in order to demonstrate asymptotic optimality via Rademacher complexity generalization techniques, and build-upon a notion called strictly local quasi-convexity in order to quantitatively characterize the optimization landscape of $\alpha $ -loss. Practically, we perform class imbalance, robustness, and classification experiments on benchmark image datasets using convolutional-neural-networks. Our main practical conclusion is that certain tasks may benefit from tuning $\alpha $ -loss away from log-loss ( $\alpha = 1$ ), and to this end we provide simple heuristics for the practitioner. In particular, navigating the $\alpha $ hyperparameter can readily provide superior model robustness to label flips ( $\alpha > 1$ ) and sensitivity to imbalanced classes ( $\alpha ).