Multi-Task and Multi-Level Detection Neural Network Based Real-Time 3D Pose Estimation

Multi-Task and Multi-Level Detection Neural Network Based Real-Time 3D Pose Estimation
复制标题

基于多任务和多级检测神经网络的实时 3D 姿态估计

DOI:
--
复制
发表时间:
2019
期刊:
Asia-Pacific Signal and Information Processing Association Annual Summit and Conference
影响因子:
--
通讯作者:
T. Ikenaga
T. Ikenaga
中科院分区:
--
文献类型:
--
作者:
Dingli Luo;Songlin Du;T. Ikenaga

文献摘要

被引文献

相似文献

三维姿态估计是人机交互和人体动作识别的核心步骤。然而,像虚拟现实这样对时间敏感的应用也需要这个任务来实现实时速度。本文提出了一种多任务、多层次的神经网络结构,具有高速友好的三维人体姿态表示。在此基础上,构建了以单幅RGB图像为输入的实时多人三维姿态估计系统。该网络通过多任务设计直接从输入图像中估计三维姿态,并通过多级检测设计保持精度和速度。通过评估,我们的系统在RTX 2080上实现了21 fps,与相关工作相比,精度仅损失33 mm。我们还提供网络可视化,以证明我们的网络工作,因为我们的设计。这项工作显示了基于单一RGB图像的3D姿态估计系统实现实时速度的可能性,这为构建低成本的3D动作捕捉系统奠定了基础。
3D pose estimation is a core step for human-computer interaction and human action recognition. However, time-sensitive applications like virtual reality also need this task to achieve real-time speed. This paper proposes a multitask and multi-level neural network architecture with a highspeed friendly 3D human pose representation. Based on this, we build a real-time multi-person 3D pose estimation system with a single RGB image as input. The network estimates 3D poses from the input image directly by the multi-task design and keeps both accuracy and speed by the multi-level detection design. By evaluation, we show our system achieves the 21 fps on RTX 2080 with only 33 mm accuracy lose compared with related works. We also provide network visualization to prove our network work as we design. This work shows the possibility for a single RGB image based 3D pose estimation system to achieve real-time speed, which is a basement for building a low-cost 3D motion capture system.