Robust Visual Tracking and Vehicle Classification via Sparse Representation

Robust Visual Tracking and Vehicle Classification via Sparse Representation
复制标题

DOI:
10.1109/tpami.2011.66
复制
发表时间:
2011-11-01
影响因子:
23.6
通讯作者:
Ling, Haibin
Ling, Haibin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mei, Xue;Ling, Haibin

文献摘要

被引文献

相似文献

在本文中,我们通过将跟踪作为粒子滤清器框架中的稀疏近似问题提出了一种强大的视觉跟踪方法。在此框架中,通过一组微不足道的模板无缝地解决了遮挡,噪声和其他具有挑战性的问题。具体而言,要在新框架中找到跟踪目标,每个目标候选者在由目标模板和琐碎模板跨越的空间中稀疏表示。通过解决L(1)规范化的最小二乘问题来实现稀疏性。然后,将最小投影误差的候选人视为跟踪目标。之后,使用贝叶斯州推理框架继续跟踪。两种策略用于进一步改善跟踪性能。首先,对目标模板进行动态更新以捕获外观变化。其次,实施非负约束以滤除杂物,而这些混乱类似于跟踪目标。我们在涉及不同类型的挑战的众多序列上测试了提出的方法,包括遮挡和照明,尺度和姿势的变化。与先前提出的跟踪器相比,提出的方法表现出了出色的性能。我们还通过引入一个静态模板集来扩展同时跟踪和识别的方法,该模板集存储来自不同类的目标图像。每个帧的识别结果都会传播,以产生整个视频的最终结果。使用室外红外视频序列在车辆跟踪和分类任务上验证了该方法。
In this paper, we propose a robust visual tracking method by casting tracking as a sparse approximation problem in a particle filter framework. In this framework, occlusion, noise, and other challenging issues are addressed seamlessly through a set of trivial templates. Specifically, to find the tracking target in a new frame, each target candidate is sparsely represented in the space spanned by target templates and trivial templates. The sparsity is achieved by solving an l(1)-regularized least-squares problem. Then, the candidate with the smallest projection error is taken as the tracking target. After that, tracking is continued using a Bayesian state inference framework. Two strategies are used to further improve the tracking performance. First, target templates are dynamically updated to capture appearance changes. Second, nonnegativity constraints are enforced to filter out clutter which negatively resembles tracking targets. We test the proposed approach on numerous sequences involving different types of challenges, including occlusion and variations in illumination, scale, and pose. The proposed approach demonstrates excellent performance in comparison with previously proposed trackers. We also extend the method for simultaneous tracking and recognition by introducing a static template set which stores target images from different classes. The recognition result at each frame is propagated to produce the final result for the whole video. The approach is validated on a vehicle tracking and classification task using outdoor infrared video sequences.