Articulated Human Detection with Flexible Mixtures of Parts

Articulated Human Detection with Flexible Mixtures of Parts
复制标题

DOI:
10.1109/tpami.2012.261
复制
发表时间:
2013-12-01
影响因子:
23.6
通讯作者:
Ramanan, Deva
Ramanan, Deva
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yang, Yi;Ramanan, Deva

文献摘要

被引文献

相似文献

我们描述了一种基于可变形部分模型的新表示的静态图像中铰接人体检测和人体姿态估计的方法。而不是使用一系列扭曲(旋转和缩短)模板建模关节,我们使用小的,非定向部分的混合物。我们描述了一个通用的、灵活的混合模型,该模型联合捕获了零件位置之间的空间关系和零件混合物之间的共现关系,增强了仅编码空间关系的标准图形结构模型。我们的模型有几个值得注意的特性:1)它们通过在相似的弯道上共享计算来有效地建模铰接,2)它们通过局部混合的组成来有效地建模一个指数级大的全局混合集合,以及3)它们捕获全局几何对局部外观的依赖(部件在不同位置看起来不同)。当关系是树状结构时,我们的模型可以有效地进行动态规划优化。我们学习所有参数,包括局部外观,空间关系和共现关系(编码局部刚性)与结构化SVM求解器。由于我们的模型足够高效,可以用作在尺度和图像位置上搜索的检测器,我们引入了新的标准来评估姿态估计和人体检测,无论是单独的还是联合的。我们表明,目前使用的评估标准可能将这两个问题混为一谈。大多数先前的方法都是用彼此独立训练的刚性和铰接模板来建模肢体,而我们提出了广泛的诊断评估,表明灵活的结构和关节训练对强表现至关重要。我们在标准基准上展示了实验结果,表明我们的方法是姿态估计的最先进系统,改进了过去在具有挑战性的Parse和Buffy数据集上的工作,同时速度提高了几个数量级。
We describe a method for articulated human detection and human pose estimation in static images based on a new representation of deformable part models. Rather than modeling articulation using a family of warped (rotated and foreshortened) templates, we use a mixture of small, nonoriented parts. We describe a general, flexible mixture model that jointly captures spatial relations between part locations and co-occurrence relations between part mixtures, augmenting standard pictorial structure models that encode just spatial relations. Our models have several notable properties: 1) They efficiently model articulation by sharing computation across similar warps, 2) they efficiently model an exponentially large set of global mixtures through composition of local mixtures, and 3) they capture the dependency of global geometry on local appearance (parts look different at different locations). When relations are tree structured, our models can be efficiently optimized with dynamic programming. We learn all parameters, including local appearances, spatial relations, and co-occurrence relations (which encode local rigidity) with a structured SVM solver. Because our model is efficient enough to be used as a detector that searches over scales and image locations, we introduce novel criteria for evaluating pose estimation and human detection, both separately and jointly. We show that currently used evaluation criteria may conflate these two issues. Most previous approaches model limbs with rigid and articulated templates that are trained independently of each other, while we present an extensive diagnostic evaluation that suggests that flexible structure and joint training are crucial for strong performance. We present experimental results on standard benchmarks that suggest our approach is the state-of-the-art system for pose estimation, improving past work on the challenging Parse and Buffy datasets while being orders of magnitude faster.