Struck: Structured Output Tracking with Kernels

Struck: Structured Output Tracking with Kernels
复制标题

DOI:
10.1109/tpami.2015.2509974
复制
发表时间:
2016-10-01
影响因子:
23.6
通讯作者:
Torr, Philip H. S.
Torr, Philip H. S.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hare, Sam;Golodetz, Stuart;Torr, Philip H. S.

文献摘要

被引文献

相似文献

自适应检测跟踪方法广泛应用于计算机视觉中对任意目标的跟踪。目前的方法将跟踪问题视为一个分类任务,并使用在线学习技术来更新对象模型。然而,为了进行这些更新,需要将估计的对象位置转换为一组标记的训练样本,并且不清楚如何最好地执行这一中间步骤。此外,分类器的目标(标签预测)没有显式地耦合到跟踪器的目标(对象位置的估计)。本文提出了一种基于结构化输出预测的自适应视觉目标跟踪框架。通过显式地允许输出空间表达跟踪器的需求,我们避免了中间分类步骤的需要。我们的方法使用核化的结构化输出支持向量机(SVM),它是在线学习的,以提供自适应跟踪。为了让我们的跟踪器在高帧速率下运行,我们(A)引入了一种预算机制,以防止在跟踪过程中可能出现的支持向量数量的无限增长,并(B)展示了如何在GPU上实现跟踪。实验表明,我们的算法在各种基准视频上的性能都优于最先进的跟踪器。此外,我们还表明,我们可以很容易地将其他功能和内核整合到我们的框架中,从而提高跟踪性能。
Adaptive tracking-by-detection methods are widely used in computer vision for tracking arbitrary objects. Current approaches treat the tracking problem as a classification task and use online learning techniques to update the object model. However, for these updates to happen one needs to convert the estimated object position into a set of labelled training examples, and it is not clear how best to perform this intermediate step. Furthermore, the objective for the classifier (label prediction) is not explicitly coupled to the objective for the tracker (estimation of object position). In this paper, we present a framework for adaptive visual object tracking based on structured output prediction. By explicitly allowing the output space to express the needs of the tracker, we avoid the need for an intermediate classification step. Our method uses a kernelised structured output support vector machine (SVM), which is learned online to provide adaptive tracking. To allow our tracker to run at high frame rates, we (a) introduce a budgeting mechanism that prevents the unbounded growth in the number of support vectors that would otherwise occur during tracking, and (b) show how to implement tracking on the GPU. Experimentally, we show that our algorithm is able to outperform state-of-the-art trackers on various benchmark videos. Additionally, we show that we can easily incorporate additional features and kernels into our framework, which results in increased tracking performance.