A Causal And-Or Graph Model for Visibility Fluent Reasoning in Tracking Interacting Objects

A Causal And-Or Graph Model for Visibility Fluent Reasoning in Tracking Interacting Objects
复制标题

DOI:
10.1109/cvpr.2018.00232
复制
发表时间:
2017-09
期刊:
2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Yuanlu Xu;Lei Qin;Xiaobai Liu;Jianwen Xie;Song-Chun Zhu
Yuanlu Xu;Lei Qin;Xiaobai Liu;Jianwen Xie;Song-Chun Zhu
中科院分区:
其他
文献类型:
--
作者:
Yuanlu Xu;Lei Qin;Xiaobai Liu;Jianwen Xie;Song-Chun Zhu

文献摘要

被引文献

相似文献

跟踪与其他对象或环境交互的人在视觉跟踪中仍然没有解决,因为视频中感兴趣的人的可见性是未知的,并且可能随时间而变化。特别是,对于最先进的人类跟踪器来说,在具有频繁人类交互的拥挤场景中恢复完整的人类轨迹仍然是困难的。在这项工作中,我们认为一个主体的可见性状态作为一个流畅的变量,其变化主要归因于主体与周围环境的相互作用,例如,引入因果与或图(CausalAnd-Or Graph,C-AOG)来表示对象的可见性流与其活动之间的因果关系,并建立概率图模型来联合推理可见性流的变化(例如,从可见到不可见)并在视频中跟踪人类。我们将这个联合任务表述为对可行因果图结构的迭代搜索,该结构能够实现快速搜索算法,例如,动态规划法我们将所提出的方法应用于具有挑战性的视频序列,以评估其能力,估计可见性,流畅的变化的主题和跟踪的兴趣随着时间的推移。比较结果表明,我们的方法优于替代跟踪器,并可以恢复完整的轨迹的人在复杂的场景与频繁的人类互动。
Tracking humans that are interacting with the other subjects or environment remains unsolved in visual tracking, because the visibility of the human of interests in videos is unknown and might vary over time. In particular, it is still difficult for state-of-the-art human trackers to recover complete human trajectories in crowded scenes with frequent human interactions. In this work, we consider the visibility status of a subject as a fluent variable, whose change is mostly attributed to the subject's interaction with the surrounding, e.g., crossing behind another object, entering a building, or getting into a vehicle, etc. We introduce a Causal And-Or Graph (C-AOG) to represent the causal-effect relations between an object's visibility fluent and its activities, and develop a probabilistic graph model to jointly reason the visibility fluent change (e.g., from visible to invisible) and track humans in videos. We formulate this joint task as an iterative search of a feasible causal graph structure that enables fast search algorithm, e.g., dynamic programming method. We apply the proposed method on challenging video sequences to evaluate its capabilities of estimating visibility fluent changes of subjects and tracking subjects of interests over time. Results with comparisons demonstrate that our method outperforms the alternative trackers and can recover complete trajectories of humans in complicated scenarios with frequent human interactions.