VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera

VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
复制标题

DOI:
10.1145/3072959.3073596
复制
发表时间:
2017-07-01
影响因子:
6.2
通讯作者:
Theobalt, Christian
Theobalt, Christian
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mehta, Dushyant;Sridhar, Srinath;Theobalt, Christian

文献摘要

被引文献

相似文献

我们提出了第一种实时方法,使用单个RGB相机以稳定、时间上一致的方式捕捉人体完整的全局3D骨骼姿态。我们的方法将一种基于卷积神经网络(CNN)的新型姿态回归器与运动学骨骼拟合相结合。我们新颖的全卷积姿态公式可实时联合回归2D和3D关节位置,并且不需要紧密裁剪的输入帧。一种实时运动学骨骼拟合方法利用CNN的输出,在连贯的运动学骨骼基础上产生时间上稳定的3D全局姿态重建。这使得我们的方法成为第一种可用于实时应用(如3D角色控制)的单目RGB方法——到目前为止,此类应用的唯一单目方法使用的是专门的RGB - D相机。我们方法的准确性在定量上与最佳的离线3D单目RGB姿态估计方法相当。我们的结果在质量上与单目RGB - D方法(如Kinect)的结果相当,有时甚至更好。然而,我们表明我们的方法比RGB - D解决方案应用更广泛,即它适用于室外场景、社区视频以及低质量的普通RGB相机。
We present the first real-time method to capture the full global 3D skeletal pose of a human in a stable, temporally consistent manner using a single RGB camera. Our method combines a new convolutional neural network (CNN) based pose regressor with kinematic skeleton fitting. Our novel fully-convolutional pose formulation regresses 2D and 3D joint positions jointly in real time and does not require tightly cropped input frames. A real-time kinematic skeleton fitting method uses the CNN output to yield temporally stable 3D global pose reconstructions on the basis of a coherent kinematic skeleton. This makes our approach the first monocular RGB method usable in real-time applications such as 3D character control-thus far, the only monocular methods for such applications employed specialized RGB-D cameras. Our method's accuracy is quantitatively on par with the best offline 3D monocular RGB pose estimation methods. Our results are qualitatively comparable to, and sometimes better than, results from monocular RGB-D approaches, such as the Kinect. However, we show that our approach is more broadly applicable than RGB-D solutions, i.e., it works for outdoor scenes, community videos, and low quality commodity RGB cameras.