Robust Bayesian Vision Transformer for Image Analysis and Classification
Robust Bayesian Vision Transformer for Image Analysis and Classification
复制标题
DOI:
10.1109/wnyispw60588.2023.10349631
复制
发表时间:
2023-11
期刊:
影响因子:
--
通讯作者:
Fazlur Rahman Bin Karim;Dimah Dera
中科院分区:
文献类型:
--
作者:
Fazlur Rahman Bin Karim;Dimah Dera
The tremendous success of the Transformer neural networks in natural language processing (NLP) boosts the interest in integrating and applying Transformer models to computer vision applications. The Vision Transformer (ViT) model adeptly captures extensive inter-dependencies among input sequences by employing the self-attention mechanism, thereby transforming picture data into semantically significant representations. In recent times, ViT has demonstrated superior performance in image classification tasks by implementing the transformer architecture, surpassing the capabilities of convolutional neural networks. Nevertheless, these deterministic architectures cannot evaluate the uncertainty associated with predictions, a crucial aspect in divergent and noisy situations. In order to guarantee the effectiveness and reliability of ViT in critical applications, the Bayesian Inference facilitates the process of making probabilistic predictions. Estimating the Bayesian posterior distribution of the network parameters enables a systematic method for reasoning about predictive uncertainty. The major difficulty in this process lies in propagating the posterior distribution through numerous non-linear layers of ViT architecture, which is mathematically cumbersome. In this paper, we propose a Bayesian Vision Transformer (Bayes-ViT) model, which seeks to make predictions as well as quantify the uncertainty associated with the output decision. The variational optimization approximates the posterior distribution over the unknown model parameters by minimizing the evidence lower bound (ELBO) loss function. The variational moments are propagated through the sequential, non-linear layers of Bayes-ViT by employing the first-order Taylor approximation. The covariance matrix of the predictive distribution effectively manifests the uncertainty associated with the output prediction. Extensive experiments on benchmark datasets (MNIST and Fashion-MNIST) exhibit (1) the superior robustness against noise and adversarial attacks compared to the deterministic ViT and (2) the self-evaluation ability based on the prediction uncertainty that becomes more evident when noise levels increase.