Grounding-Tracking-Integration

Grounding-Tracking-Integration
复制标题

DOI:
10.1109/tcsvt.2020.3038720
复制
发表时间:
2019-12
影响因子:
8.4
通讯作者:
Zhengyuan Yang;T. Kumar;Tianlang Chen;Jiebo Luo
Zhengyuan Yang;T. Kumar;Tianlang Chen;Jiebo Luo
中科院分区:
工程技术1区
文献类型:
--
作者:
Zhengyuan Yang;T. Kumar;Tianlang Chen;Jiebo Luo

文献摘要

被引文献

相似文献

本文研究了基于语言查询的视频目标框序列的语言跟踪。我们提出了一个名为GTI的框架,它将问题分解为三个子任务:接地、跟踪和集成。三个子任务模块同时工作,逐帧预测盒序列。“接地”直接从语言查询中预测所引用的区域。“跟踪”基于前一帧中接地区域的历史来定位目标。“整合”通过协同结合接地和跟踪产生最终预测。我们以“整合”任务为重点,探索如何在每一帧中显示接地区域的质量,并实现期望的互利组合。为此,我们提出了一种“RT-integration”方法,通过定义和预测两个分数来指导整合:1)R-score代表区域的正确性,即接地预测是否准确地覆盖了目标;2)T-score代表模板质量,即该区域是否提供了信息丰富的视觉线索,以改善未来帧的跟踪。我们提出了我们的实时GTI实现与建议的rt集成,并在LaSOT和语种OTB99上对框架进行了基准测试,结果非常有希望。此外,我们还制作了LaSOT查询的消歧版本,以方便语言研究的未来跟踪。
In this paper, we study tracking by language that localizes the target box sequence in a video based on a language query. We propose a framework called GTI that decomposes the problem into three sub-tasks: Grounding, Tracking, and Integration. The three sub-task modules operate simultaneously and predict the box sequence frame-by-frame. “Grounding” predicts the referred region directly from the language query. “Tracking” localizes the target based on the history of the grounded regions in previous frames. “Integration” generates final predictions by synergistically combining grounding and tracking. With the “integration” task as the key, we explore how to indicate the quality of the grounded regions in each frame and achieve the desired mutually beneficial combination. To this end, we propose an “RT-integration” method that defines and predicts two scores to guide the integration: 1) R-score represents the Region correctness whether the grounding prediction accurately covers the target, and 2) T-score represents the Template quality whether the region provides informative visual cues to improve tracking in future frames. We present our real-time GTI implementation with the proposed RT-integration, and benchmark the framework on LaSOT and Lingual OTB99 with highly promising results. Moreover, we produce a disambiguated version of LaSOT queries to facilitate future tracking by language studies.