Single Image Depth Map Estimation for Improving Posture Recognition

Single Image Depth Map Estimation for Improving Posture Recognition
复制标题

用于改进姿势识别的单图像深度图估计

DOI:
10.1109/jsen.2021.3122128
复制
发表时间:
2021-12-01
影响因子:
4.3
通讯作者:
Chen, Yen-Wei
Chen, Yen-Wei
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Liu, Jiaqing;Tsujinaga, Seiju;Chen, Yen-Wei

文献摘要

被引文献

相似文献

基于图像的姿态识别是一个非常具有挑战性的问题,因为很难从彩色图像中获取丰富的三维信息。为了解决这个问题,我们提出了一个新的和统一的框架,人体姿势识别,应用单图像深度图估计彩色图像。所提出的方法包括两个阶段。第一阶段通过改进的Pix2Pix生成模块从单色图像中估计深度图。生成模块配备了一个混合损失函数,它捕获高级特征并恢复尖锐的深度不连续性,从而改善深度估计结果。第二阶段(识别阶段)通过结合估计的深度图来提高基于彩色图像的识别性能。因此,开发了一种双流CNN架构,该架构分别处理彩色图像及其估计的深度图像,用于鲁棒的姿态识别。为了验证其有效性,我们首先测试所提出的方法在一个新的姿势数据集,其中包含13800个样本的配对颜色和深度的6个主题与15个姿势。本工作中使用的数据集已创建并发布,可在http://media.ritsumei.ac.jp/iipl/database/pose/上获得。在公开的OUHANDS手势数据集上进行了大量的实验。实验结果表明,该方法在人体姿态和手势识别任务上都取得了上级的性能。
Image-based posture recognition is a very challenging problem since it is difficult to acquire rich 3D information from the posture in color image. To address this issue, we present a novel and unified framework for human posture recognition, applying single image depth map estimation from color images. The proposed method includes two stages. The first stage estimates the depth map from the single-color image by an improved Pix2Pix generation module. The generation module is equipped with a hybrid loss function that captures the high-level features and recovers the sharp depth discontinuities, thus improving the depth estimation results. The second stage (the recognition stage) improves the color image-based recognition performance by incorporating the estimated depth map. Thereby, a two-stream CNN architecture that separately processes the color image and its estimated depth image is developed for robust posture recognition. To verify its effectiveness, we first test the proposed method on a novel pose dataset, which contains 13800 samples of paired color-and-depth of 6 subjects with 15 poses. The dataset used in this work is been created and released, is available at http://media.ritsumei.ac.jp/iipl/database/pose/. Extensive experiments are also performed on the public OUHANDS hand gesture dataset. Experiments demonstrate that the proposed method achieves superior performance on both human pose and hand gesture recognition tasks.