A Survey on Vision Transformer

A Survey on Vision Transformer
复制标题

视觉Transformer研究综述

DOI:
10.1109/tpami.2022.3152247
复制
发表时间:
2023-01-01
影响因子:
23.6
通讯作者:
Tao, Dacheng
Tao, Dacheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Han, Kai;Wang, Yunhe;Tao, Dacheng

文献摘要

被引文献

相似文献

Transformer是一种主要基于自注意机制的深度神经网络,最早应用于自然语言处理领域。由于其强大的表示能力,研究人员正在寻找将变压器应用于计算机视觉任务的方法。在各种视觉基准测试中,基于变压器的模型的表现与其他类型的网络(如卷积和循环神经网络)相似或更好。变压器以其高性能和对视觉特定感应偏置的需求小而受到计算机视觉界越来越多的关注。本文对这些视觉转换模型进行了综述,并对它们在不同任务中的应用进行了分类,分析了它们的优缺点。我们探索的主要类别包括骨干网、高级/中级视觉、低级视觉和视频处理。我们还包括有效的变压器方法,用于将变压器推向实际的基于设备的应用程序。此外,我们还简要介绍了计算机视觉中的自注意机制,因为它是变压器的基本组成部分。在文章的最后,我们讨论了视觉变压器面临的挑战,并提出了未来的研究方向。
Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong representation capabilities, researchers are looking at ways to apply transformer to computer vision tasks. In a variety of visual benchmarks, transformer-based models perform similar to or better than other types of networks such as convolutional and recurrent neural networks. Given its high performance and less need for vision-specific inductive bias, transformer is receiving more and more attention from the computer vision community. In this paper, we review these vision transformer models by categorizing them in different tasks and analyzing their advantages and disadvantages. The main categories we explore include the backbone network, high/mid-level vision, low-level vision, and video processing. We also include efficient transformer methods for pushing transformer into real device-based applications. Furthermore, we also take a brief look at the self-attention mechanism in computer vision, as it is the base component in transformer. Toward the end of this paper, we discuss the challenges and provide several further research directions for vision transformers.