How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age Estimation

How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age Estimation
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
A. Akbari;Muhammad Awais;M. Bashar;J. Kittler
A. Akbari;Muhammad Awais;M. Bashar;J. Kittler
中科院分区:
其他
文献类型:
--
作者:
A. Akbari;Muhammad Awais;M. Bashar;J. Kittler

文献摘要

相似文献

在由许多外部和内部因素引起的各种领域中具有良好的泛化性能是任何机器学习算法的基本目标。本文从理论上证明了损失函数的选择对于提高基于深度学习的系统的泛化性能至关重要。通过推导由随机梯度下降训练的深度神经模型的泛化误差界,我们确定了与泛化误差相关的损失函数的特征,因此可以用于指导损失函数的选择过程。综上所述,本文的主要论述是:选择稳定的损失函数,推广效果更好。针对计算机视觉中一个具有挑战性的课题--从人脸中估计人的年龄,我们提出了一种新的损失函数来解决这个学习问题。我们从理论上证明,建议的损失函数实现更强的稳定性,从而更严格的推广误差界,相比其他常见的损失函数为这个问题。我们在理论上支持了我们的发现,并在实验上证明了指导过程的优点,取得了重大改进。
Good generalization performance across a wide variety of domains caused by many external and internal factors is the fundamental goal of any machine learning algorithm. This paper theoretically proves that the choice of loss function matters for improving the generalization performance of deep learning-based systems. By deriving the generalization error bound for deep neural models trained by stochastic gradient descent, we pinpoint the characteristics of the loss function that is linked to the generalization error, and can therefore be used for guiding the loss function selection process. In summary, our main statement in this paper is: choose a stable loss function, generalize better. Focusing on human age estimation from the face which is a challenging topic in computer vision, we then propose a novel loss function for this learning problem. We theoretically prove that the proposed loss function achieves stronger stability, and consequently a tighter generalization error bound, compared to the other common loss functions for this problem. We have supported our findings theoretically, and demonstrated the merits of the guidance process experimentally, achieving significant improvements.