Learning a Deep Compact Image Representation for Visual Tracking

Learning a Deep Compact Image Representation for Visual Tracking
复制标题

DOI:
--
复制
发表时间:
2013-12
期刊:
--
影响因子:
--
通讯作者:
Naiyan Wang;D. Yeung
Naiyan Wang;D. Yeung
中科院分区:
其他
文献类型:
--
作者:
Naiyan Wang;D. Yeung

文献摘要

被引文献

相似文献

在本文中,我们研究了具有挑战性的问题,跟踪运动对象的轨迹在视频可能非常复杂的背景。与大多数现有的跟踪器相比,它们只在线学习被跟踪对象的外观,我们采取了不同的方法,受到深度学习架构最新进展的启发,更多地强调(无监督)特征学习问题。具体来说,通过使用辅助自然图像,我们离线训练了一个堆栈去噪自动编码器,以学习对变化更鲁棒的通用图像特征。然后,知识从离线培训转移到在线跟踪过程。在线跟踪涉及一个分类神经网络,它是从训练的自动编码器的编码器部分作为一个特征提取器和一个额外的分类层。特征提取器和分类器都可以进一步调整以适应移动对象的外观变化。在一些具有挑战性的基准视频序列上与最先进的跟踪器进行比较表明,当我们的跟踪器的MATLAB实现与适度的图形处理单元(GPU)一起使用时,我们的深度学习跟踪器更准确,同时保持低计算成本和实时性能。
In this paper, we study the challenging problem of tracking the trajectory of a moving object in a video with possibly very complex background. In contrast to most existing trackers which only learn the appearance of the tracked object online, we take a different approach, inspired by recent advances in deep learning architectures, by putting more emphasis on the (unsupervised) feature learning problem. Specifically, by using auxiliary natural images, we train a stacked de-noising autoencoder offline to learn generic image features that are more robust against variations. This is then followed by knowledge transfer from offline training to the online tracking process. Online tracking involves a classification neural network which is constructed from the encoder part of the trained autoencoder as a feature extractor and an additional classification layer. Both the feature extractor and the classifier can be further tuned to adapt to appearance changes of the moving object. Comparison with the state-of-the-art trackers on some challenging benchmark video sequences shows that our deep learning tracker is more accurate while maintaining low computational cost with real-time performance when our MATLAB implementation of the tracker is used with a modest graphics processing unit (GPU).