Toward Accurate Pixelwise Object Tracking via Attention Retrieval

Toward Accurate Pixelwise Object Tracking via Attention Retrieval
复制标题

DOI:
10.1109/tip.2021.3117077
复制
发表时间:
2020-08
影响因子:
10.6
通讯作者:
Zhipeng Zhang;Yufan Liu;Bing Li;Weiming Hu;Houwen Peng
Zhipeng Zhang;Yufan Liu;Bing Li;Weiming Hu;Houwen Peng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhipeng Zhang;Yufan Liu;Bing Li;Weiming Hu;Houwen Peng

文献摘要

相似文献

由于运行速度和分割精度的竞争,像素级单目标跟踪具有挑战性。当前最先进的实时方法通过共享骨干网络的计算来无缝地连接跟踪和分段,例如,SiamMask和D3 S从跟踪模型分叉一个轻分支来预测分割掩码。虽然有效,但直接重用来自跟踪网络的特征可能会损害分割精度,因为主干特征中的背景杂波往往会在分割中引入误报。为了缓解这个问题,我们提出了一个统一的跟踪检索分割框架组成的注意力检索网络(ARN)和迭代反馈网络(IFN)。而不是分割的目标内的包围盒,所提出的框架进行软空间约束的骨干功能,以获得一个准确的全球分割地图。具体地,在ARN中,首先通过充分使用第一帧的信息来构建查找表(LUT)。通过检索它,目标感知的注意力地图生成抑制背景杂波的负面影响。为了进一步细化分割的轮廓,IFN通过将预测的掩模作为反馈指导来迭代增强不同分辨率的特征。我们的框架在最近的像素跟踪基准VOT 2020上设置了一个新的艺术状态,并以40 fps运行。值得注意的是,该模型在VOT 2020、DAVIS 2016和DAVIS 2017上分别超过SiamMask 11.7/4.2/5.5点。代码可在https://github.com/JudasDie/SOTS上获得。
Pixelwise single object tracking is challenging due to the competition of running speeds and segmentation accuracy. Current state-of-the-art real-time approaches seamlessly connect tracking and segmentation by sharing computation of the backbone network, e.g., SiamMask and D3S fork a light branch from the tracking model to predict segmentation mask. Although efficient, directly reusing features from tracking networks may harm the segmentation accuracy, since background clutter in the backbone feature tends to introduce false positives in segmentation. To mitigate this problem, we propose a unified tracking-retrieval-segmentation framework consisting of an attention retrieval network (ARN) and an iterative feedback network (IFN). Instead of segmenting the target inside the bounding box, the proposed framework performs soft spatial constraints on backbone features to obtain an accurate global segmentation map. Concretely, in ARN, a look-up-table (LUT) is first built by sufficiently using the information of the first frame. By retrieving it, a target-aware attention map is generated to suppress the negative influence of background clutter. To ulteriorly refine the contour of the segmentation, IFN iteratively enhances the features at different resolutions by taking the predicted mask as feedback guidance. Our framework sets a new state of the art on the recent pixelwise tracking benchmark VOT2020 and runs at 40 fps. Notably, the proposed model surpasses SiamMask by 11.7/4.2/5.5 points on VOT2020, DAVIS2016, and DAVIS2017, respectively. Code is available at https://github.com/JudasDie/SOTS.