Multi-Scale Structure-Aware Network for Human Pose Estimation

Multi-Scale Structure-Aware Network for Human Pose Estimation
复制标题

DOI:
10.1007/978-3-030-01216-8_44
复制
发表时间:
2018-03
期刊:
--
影响因子:
--
通讯作者:
Lipeng Ke;Ming-Ching Chang;H. Qi;Siwei Lyu
Lipeng Ke;Ming-Ching Chang;H. Qi;Siwei Lyu
中科院分区:
其他
文献类型:
--
作者:
Lipeng Ke;Ming-Ching Chang;H. Qi;Siwei Lyu

文献摘要

被引文献

相似文献

我们开发了一种鲁棒的多尺度结构感知神经网络用于人体姿态估计。该方法对现有的深度卷积-反卷积沙漏模型进行了改进,主要有四个方面的改进:(1)多尺度监督,通过结合跨尺度特征热图来加强匹配主体关键点的上下文特征学习;(2)多尺度回归网络在最后对多尺度特征的结构匹配进行全局优化;(3)在中间监督和回归中使用结构感知损失来改善关键点与各自相邻点的匹配(4)一个关键点掩蔽训练方案,该方案可以有效地微调我们的网络,通过相邻的匹配来鲁棒地定位被遮挡的关键点。我们的方法可以有效地改进目前最先进的姿态估计方法,这些方法在尺度变化、遮挡和复杂的多人场景中存在困难。这种多尺度监督与回归网络紧密集成,可以有效地(i)使用多尺度特征集合来定位关键点,(ii)通过最大化多个关键点和尺度的结构一致性来推断全局姿态配置。关键点掩蔽训练增强了这些优点,使学习集中在硬遮挡样本上。我们的方法在最先进的方法中在MPII挑战排行榜上处于领先地位。
We develop a robust multi-scale structure-aware neural network for human pose estimation. This method improves the recent deep conv-deconv hourglass models with four key improvements:(1) multi-scale supervision to strengthen contextual feature learning in matching body keypoints by combining feature heatmaps across scales,(2) multi-scale regression network at the end to globally optimize the structural matching of the multi-scale features,(3) structure-aware loss used in the intermediate supervision and at the regression to improve the matching of keypoints and respective neighbors to infer a higher-order matching configurations, and (4) a keypoint masking training scheme that can effectively fine-tune our network to robustly localize occluded keypoints via adjacent matches. Our method can effectively improve state-of-the-art pose estimation methods that suffer from difficulties in scale varieties, occlusions, and complex multi-person scenarios. This multi-scale supervision tightly integrates with the regression network to effectively (i) localize keypoints using the ensemble of multi-scale features, and (ii) infer global pose configuration by maximizing structural consistencies across multiple keypoints and scales. The keypoint masking training enhances these advantages to focus learning on hard occlusion samples. Our method achieves the leading position in the MPII challenge leaderboard among the state-of-the-art methods.