Calibration-Aware Bayesian Learning

Calibration-Aware Bayesian Learning
复制标题

DOI:
10.1109/mlsp55844.2023.10285894
复制
发表时间:
2023-05
期刊:
2023 IEEE 33rd International Workshop on Machine Learning for Signal Processing (MLSP)
影响因子:
--
通讯作者:
Jiayi Huang;Sangwoo Park;O. Simeone
Jiayi Huang;Sangwoo Park;O. Simeone
中科院分区:
其他
文献类型:
--
作者:
Jiayi Huang;Sangwoo Park;O. Simeone

文献摘要

相似文献

众所周知,深度学习模型,包括像大型语言模型这样的现代系统,对其决策的不确定性提供了不可靠的估计。为了改善模型的置信度的质量,也称为校准,通常的方法需要在训练损失中添加依赖于数据或独立于数据的正则化项。依赖于数据的正则化方法最近被引入到传统频率学习的背景下,以惩罚置信度和准确度之间的偏差。相反,与数据无关的正则化是贝叶斯学习的核心,强制模型参数空间中的变分分布遵守先验密度。前一种方法无法量化认知不确定性,而后一种方法受到模型错误说明的严重影响。鉴于这两种方法的局限性,本文提出了一种集成的框架,称为校准感知贝叶斯神经网络(CA-BNN),它应用了两种正则化方法,同时像贝叶斯学习一样在变分分布上进行优化。数值结果验证了该方法在期望校准误差和可靠性图方面的优势。
Deep learning models, including modern systems like large language models, are well known to offer unreliable estimates of the uncertainty of their decisions. In order to improve the quality of the confidence levels, also known as calibration, of a model, common approaches entail the addition of either data-dependent or data-independent regularization terms to the training loss. Data-dependent regularizers have been recently introduced in the context of conventional frequentist learning to penalize deviations between confidence and accuracy. In contrast, data-independent regularizers are at the core of Bayesian learning, enforcing adherence of the variational distribution in the model parameter space to a prior density. The former approach is unable to quantify epistemic uncertainty, while the latter is severely affected by model misspecification. In light of the limitations of both methods, this paper proposes an integrated framework, referred to as calibration-aware Bayesian neural networks (CA-BNNs), that applies both regularizers while optimizing over a variational distribution as in Bayesian learning. Numerical results validate the advantages of the proposed approach in terms of expected calibration error (ECE) and reliability diagrams.