Robust Online Tracking via Contrastive Spatio-Temporal Aware Network

Robust Online Tracking via Contrastive Spatio-Temporal Aware Network
复制标题

通过对比时空感知网络进行稳健的在线跟踪

DOI:
10.1109/tip.2021.3050314
复制
发表时间:
2021-01
影响因子:
10.6
通讯作者:
Cao Xiaochun
Cao Xiaochun
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yao Siyuan;Zhang Hua;Ren Wenqi;Ma Chao;Han Xiaoguang;Cao Xiaochun

文献摘要

参考文献

相似文献

现有的基于深度特征的检测跟踪方法近年来取得了可喜的成果。然而,这些方法主要利用从单个静态帧中学习到的特征表示,因此很少关注帧之间的时间平滑性。这很容易导致跟踪器在出现大的外观变化和遮挡时漂移。为了解决这个问题,我们提出了一个双流网络来学习判别时空特征表征来表示目标物体。该网络由空间卷积神经网络模块和时间卷积神经网络模块组成。具体来说,Spatial ConvNet采用二维卷积编码静态帧中特定目标的外观,而Temporal ConvNet使用三维卷积建模时间外观变化,并在短视频片段中学习一致的时间模式。在此基础上,提出了一种建议细化模块,对预测的边界框进行调整,使目标定位输出在视频序列中更加一致。此外,为了提高在线更新过程中模型的适应性,我们提出了一种对比在线硬样本挖掘(OHEM)策略,该策略选择硬负样本并强制将其嵌入到更具判别性的特征空间中。在OTB, Temple Color和VOT基准测试上进行的大量实验表明,所提出的算法优于最先进的方法。
Existing tracking-by-detection approaches using deep features have achieved promising results in recent years. However, these methods mainly exploit feature representations learned from individual static frames, thus paying little attention to the temporal smoothness between frames. This easily leads trackers to drift in the presence of large appearance variations and occlusions. To address this issue, we propose a two-stream network to learn discriminative spatio-temporal feature representations to represent the target objects. The proposed network consists of a Spatial ConvNet module and a Temporal ConvNet module. Specifically, the Spatial ConvNet adopts 2D convolutions to encode the target-specific appearance in static frames, while the Temporal ConvNet models the temporal appearance variations using 3D convolutions and learns consistent temporal patterns in a short video clip. Then we propose a proposal refinement module to adjust the predicted bounding box, which can make the target localizing outputs to be more consistent in video sequences. In addition, to improve the model adaptation during online update, we propose a contrastive online hard example mining (OHEM) strategy, which selects hard negative samples and enforces them to be embedded in a more discriminative feature space. Extensive experiments conducted on the OTB, Temple Color and VOT benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods.
在全球范围内高效地挖掘硬样本以进行人员重新识别
DOI: 10.1109/jiot.2020.2980549
发表时间: 2020-10
影响因子: 10.6
作者:
Hao Sheng;Yanwei Zheng;Wei Ke;Dongxiao Yu;Xiuzhen Cheng;Weifeng Lyu;Zhang Xiong
通讯作者: Zhang Xiong
DOI: --
发表时间: 2015-02
期刊: --
影响因子: --
作者:
Seunghoon Hong;Tackgeun You;Suha Kwak;Bohyung Han
通讯作者: Seunghoon Hong;Tackgeun You;Suha Kwak;Bohyung Han
DOI: 10.1109/tpami.2018.2858826
发表时间: 2020-02-01
影响因子: 23.6
作者:
Lin, Tsung-Yi;Goyal, Priya;Dollar, Piotr
通讯作者: Dollar, Piotr
DOI: 10.1016/j.neucom.2016.08.070
发表时间: 2016-07
期刊: ArXiv
影响因子: --
作者:
Bohan Zhuang;Lijun Wang;Huchuan Lu
通讯作者: Bohan Zhuang;Lijun Wang;Huchuan Lu
DOI: --
发表时间: --
期刊: --
影响因子: --
作者:
João F. Henriques;Rui Caseiro;P. Martins;Jorge Batista
通讯作者: João F. Henriques;Rui Caseiro;P. Martins;Jorge Batista