Structured learning of human interactions in TV shows.

Structured learning of human interactions in TV shows.
复制标题

电视节目中人际互动的结构化学习。

DOI:
10.1109/tpami.2012.24
复制
发表时间:
2012
影响因子:
23.6
通讯作者:
Patron-Perez A
Patron-Perez A
中科院分区:
计算机科学1区
文献类型:
--
作者:
Patron-Perez A

文献摘要

相似文献

这项工作的目标是识别和时空定位的两个人的互动视频。我们的方法是以人为中心。作为第一阶段,我们使用检测跟踪方法跟踪视频中的所有上半身和头部,该方法将检测与KLT跟踪和集团分割以及遮挡检测相结合,以产生鲁棒的人跟踪。我们开发了基于头部方向(使用一组姿势特定的分类器估计)和它们周围的局部时空区域的活动的局部描述符,以及将人的相对位置编码为交互类型的函数的全局描述符。对模型的学习和推理使用结构化输出SVM,该SVM以原则性的方式结合了局部和全局描述符。使用该模型的推理产生关于哪些人对正在交互、他们的交互类和他们的头部方向的信息(这也被视为变量,使得分类器中的错误能够使用全局上下文来纠正)。我们表明,推理可以进行多项式复杂的人数,并描述了一个有效的算法。该方法在包括从23个不同电视节目中获取的300个视频剪辑的新数据集和基准UT-交互数据集上进行评估。
The objective of this work is recognition and spatiotemporal localization of two-person interactions in video. Our approach is person-centric. As a first stage we track all upper bodies and heads in a video using a tracking-by-detection approach that combines detections with KLT tracking and clique partitioning, together with occlusion detection, to yield robust person tracks. We develop local descriptors of activity based on the head orientation (estimated using a set of pose-specific classifiers) and the local spatiotemporal region around them, together with global descriptors that encode the relative positions of people as a function of interaction type. Learning and inference on the model uses a structured output SVM which combines the local and global descriptors in a principled manner. Inference using the model yields information about which pairs of people are interacting, their interaction class, and their head orientation (which is also treated as a variable, enabling mistakes in the classifier to be corrected using global context). We show that inference can be carried out with polynomial complexity in the number of people, and describe an efficient algorithm for this. The method is evaluated on a new dataset comprising 300 video clips acquired from 23 different TV shows and on the benchmark UT--Interaction dataset.