Robust Bayesian Vision Transformer for Image Analysis and Classification

Robust Bayesian Vision Transformer for Image Analysis and Classification
复制标题

DOI:
10.1109/wnyispw60588.2023.10349631
复制
发表时间:
2023-11
期刊:
2023 IEEE Western New York Image and Signal Processing Workshop (WNYISPW)
影响因子:
--
通讯作者:
Fazlur Rahman Bin Karim;Dimah Dera
Fazlur Rahman Bin Karim;Dimah Dera
中科院分区:
其他
文献类型:
--
作者:
Fazlur Rahman Bin Karim;Dimah Dera

文献摘要

相似文献

Transformer神经网络在自然语言处理(NLP)中的巨大成功,提高了将Transformer模型集成和应用于计算机视觉应用的兴趣。视觉Transformer(ViT)模型通过采用自注意机制巧妙地捕获输入序列之间的广泛相互依赖性,从而将图片数据转换为语义上重要的表示。最近,ViT通过实现Transformer架构,在图像分类任务中表现出上级性能,超越了卷积神经网络的能力。然而,这些确定性架构无法评估与预测相关的不确定性,这是发散和嘈杂情况下的一个关键方面。为了保证ViT在关键应用中的有效性和可靠性,贝叶斯推理简化了概率预测的过程。估计贝叶斯后验分布的网络参数,使一个系统的方法推理预测的不确定性。这个过程中的主要困难在于通过ViT架构的许多非线性层传播后验分布,这在数学上是麻烦的。在本文中,我们提出了一个贝叶斯视觉Transformer(贝叶斯-ViT)模型,该模型旨在进行预测以及量化与输出决策相关的不确定性。变分优化通过最小化证据下限(ELBO)损失函数来近似未知模型参数的后验分布。通过采用一阶泰勒近似的贝叶斯-ViT的顺序,非线性层的变分矩传播。预测分布的协方差矩阵有效地体现了与输出预测相关的不确定性。在基准数据集(MNIST和Fashion-MNIST)上的大量实验显示:(1)与确定性ViT相比,对噪声和对抗性攻击具有上级鲁棒性;(2)基于预测不确定性的自我评估能力,当噪声水平增加时,这种能力变得更加明显。
The tremendous success of the Transformer neural networks in natural language processing (NLP) boosts the interest in integrating and applying Transformer models to computer vision applications. The Vision Transformer (ViT) model adeptly captures extensive inter-dependencies among input sequences by employing the self-attention mechanism, thereby transforming picture data into semantically significant representations. In recent times, ViT has demonstrated superior performance in image classification tasks by implementing the transformer architecture, surpassing the capabilities of convolutional neural networks. Nevertheless, these deterministic architectures cannot evaluate the uncertainty associated with predictions, a crucial aspect in divergent and noisy situations. In order to guarantee the effectiveness and reliability of ViT in critical applications, the Bayesian Inference facilitates the process of making probabilistic predictions. Estimating the Bayesian posterior distribution of the network parameters enables a systematic method for reasoning about predictive uncertainty. The major difficulty in this process lies in propagating the posterior distribution through numerous non-linear layers of ViT architecture, which is mathematically cumbersome. In this paper, we propose a Bayesian Vision Transformer (Bayes-ViT) model, which seeks to make predictions as well as quantify the uncertainty associated with the output decision. The variational optimization approximates the posterior distribution over the unknown model parameters by minimizing the evidence lower bound (ELBO) loss function. The variational moments are propagated through the sequential, non-linear layers of Bayes-ViT by employing the first-order Taylor approximation. The covariance matrix of the predictive distribution effectively manifests the uncertainty associated with the output prediction. Extensive experiments on benchmark datasets (MNIST and Fashion-MNIST) exhibit (1) the superior robustness against noise and adversarial attacks compared to the deterministic ViT and (2) the self-evaluation ability based on the prediction uncertainty that becomes more evident when noise levels increase.