Learning Latent Representations of 3D Human Pose with Deep Neural Networks

Learning Latent Representations of 3D Human Pose with Deep Neural Networks
复制标题

DOI:
10.1007/s11263-018-1066-6
复制
发表时间:
2018-12-01
影响因子:
19.5
通讯作者:
Fua, Pascal
Fua, Pascal
中科院分区:
计算机科学2区
文献类型:
--
作者:
Katircioglu, Isinsu;Tekin, Bugra;Fua, Pascal

文献摘要

被引文献

相似文献

最近的单目3D姿态估计方法依赖于深度学习。他们要么训练卷积神经网络直接从图像回归到3D姿势,忽略人体关节之间的依赖关系,要么通过最大余量结构化学习框架对这些依赖关系进行建模,这在推理时涉及很高的计算成本。在本文中,我们介绍了一种深度学习回归架构,用于从单目图像或2D关节位置热图中结构化预测3D人体姿势,该架构依赖于过完备自动编码器来学习高维潜在姿势表示并考虑关节依赖性。我们进一步提出了一个有效的长短期记忆网络,以加强三维姿态预测的时间一致性。我们证明了我们的方法在标准3D人体姿态估计基准的结构保留和预测精度方面都达到了最先进的性能。
Most recent approaches to monocular 3D pose estimation rely on Deep Learning. They either train a Convolutional Neural Network to directly regress from an image to a 3D pose, which ignores the dependencies between human joints, or model these dependencies via a max-margin structured learning framework, which involves a high computational cost at inference time. In this paper, we introduce a Deep Learning regression architecture for structured prediction of 3D human pose from monocular images or 2D joint location heatmaps that relies on an overcomplete autoencoder to learn a high-dimensional latent pose representation and accounts for joint dependencies. We further propose an efficient Long Short-Term Memory network to enforce temporal consistency on 3D pose predictions. We demonstrate that our approach achieves state-of-the-art performance both in terms of structure preservation and prediction accuracy on standard 3D human pose estimation benchmarks.