Talking Heads: Detecting Humans and Recognizing Their Interactions

Talking Heads: Detecting Humans and Recognizing Their Interactions
复制标题

DOI:
10.1109/cvpr.2014.117
复制
发表时间:
2014-06
期刊:
2014 IEEE Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Minh Hoai;Andrew Zisserman
Minh Hoai;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Minh Hoai;Andrew Zisserman

文献摘要

被引文献

相似文献

这项工作的目的是准确和有效地检测配置的一个或多个人在编辑的电视材料。由于电影风格,这种配置通常出现在标准安排中,我们利用这一点来提供场景背景。我们做出以下贡献:首先,我们介绍了一种新的可学习的上下文感知配置模型,用于检测电视材料中的人的集合,该模型预测配置中每个上身的比例和位置,第二,我们表明可以使用动态规划全局有效地解决模型的推理,并实现最大余量学习框架,第三,我们表明,该配置模型在预测视频帧中的上身位置方面大大优于可变形部分模型(Deformable Part Model,简称DMM),即使DMM配备有其他上身的上下文。实验是在两个数据集上进行的:电视人机交互数据集,以及来自四个不同电视节目的150集。我们还证明了该模型在识别电视节目中的互动的好处。
The objective of this work is to accurately and efficiently detect configurations of one or more people in edited TV material. Such configurations often appear in standard arrangements due to cinematic style, and we take advantage of this to provide scene context. We make the following contributions: first, we introduce a new learnable context aware configuration model for detecting sets of people in TV material that predicts the scale and location of each upper body in the configuration, second, we show that inference of the model can be solved globally and efficiently using dynamic programming, and implement a maximum margin learning framework, and third, we show that the configuration model substantially outperforms a Deformable Part Model (DPM) for predicting upper body locations in video frames, even when the DPM is equipped with the context of other upper bodies. Experiments are performed over two datasets: the TV Human Interaction dataset, and 150 episodes from four different TV shows. We also demonstrate the benefits of the model in recognizing interactions in TV shows.