Exploiting Multi-View Part-Wise Correlation via an Efficient Transformer for Vehicle Re-Identification

Exploiting Multi-View Part-Wise Correlation via an Efficient Transformer for Vehicle Re-Identification
复制标题

DOI:
10.1109/tmm.2021.3134839
复制
发表时间:
2023-01-01
影响因子:
7.3
通讯作者:
Zhang, Ziming
Zhang, Ziming
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Ming;Liu, Jun;Zhang, Ziming

文献摘要

被引文献

相似文献

基于图像的车辆再识别(ReID)技术近年来取得了很大的进展。然而,大多数现有的工作努力从一个单一的图像中提取鲁棒的,但歧视性的特征来表示一个车辆实例。我们认为,从不同的观点,例如,正面和背面具有显著不同的外观和识别模式。为了识别每辆车,这些模型必须从完全不同的角度捕获一致的“ID代码”,这导致了学习困难。此外,我们主张视图之间的部分级对应,即,从相同图像观察到的各种车辆部件以及从不同视点可见的相同部件也有助于实例级特征学习。出于这些动机,我们建议通过建模部分相关性从多个视图中提取全面的车辆实例表示。为此,我们提出了我们高效的基于变换的框架,以利用车辆ReID的内部和视图间相关性。具体来说,我们首先采用一个convnet编码器来压缩一系列的补丁嵌入从每个视图。然后,我们的高效的Transformer,除了一个常规的分类令牌蒸馏令牌和噪声令牌组成,构造强制执行这些补丁嵌入相互作用,无论他们是从相同或不同的意见。我们在广泛使用的车辆ReID基准上进行了大量的实验,我们的方法达到了最先进的性能,显示了我们的方法的有效性。
Image-based vehicle re-identification (ReID) has witnessed much progress in recent years. However, most of existing works struggled to extract robust but discriminative features from a single image to represent one vehicle instance. We argue that images taken from distinct viewpoints, e.g., front and back, have significantly different appearances and patterns for recognition. In order to identify each vehicle, these models have to capture consistent "ID codes " from totally different views, causing learning difficulties. Additionally, we claim that part-level correspondences among views, i.e., various vehicle parts observed from the identical image and the same part visible from different viewpoints, contribute to instance-level feature learning as well. Motivated by these, we propose to extract comprehensive vehicle instance representations from multiple views through modelling part-wise correlations. To this end, we present our efficient transformer-based framework to exploit both inner- and inter-view correlations for vehicle ReID. In specific, we first adopt a convnet encoder to condense a series of patch embeddings from each view. Then our efficient transformer, consisting of a distillation token and a noise token in addition to a regular classification token, is constructed for enforcing these patch embeddings to interact with each other regardless of whether they are taken from identical or different views. We conduct extensive experiments on widely used vehicle ReID benchmarks, and our approach achieves the state-of-the-art performance, showing the effectiveness of our method.