Fusing Visual and Inertial Sensors with Semantics for 3D Human Pose Estimation

Fusing Visual and Inertial Sensors with Semantics for 3D Human Pose Estimation
复制标题

DOI:
10.1007/s11263-018-1118-y
复制
发表时间:
2019-04-01
影响因子:
19.5
通讯作者:
Collomosse, John
Collomosse, John
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gilbert, Andrew;Trumble, Matthew;Collomosse, John

文献摘要

被引文献

相似文献

我们提出了一种方法来准确地估计三维人体姿态融合多视点视频(MVV)与惯性测量单元(IMU)传感器数据,没有光学标记,复杂的硬件设置或全身模型。独特地,我们使用多通道3D卷积神经网络来学习从离散化体积概率视觉船体中的MVV的视觉占用和语义2D姿态估计的姿态嵌入。学习的姿势流与IMU数据的正向运动学求解并行处理,时间模型(LSTM)利用求解的关节之间丰富的空间和时间长范围依赖性,然后将两个流融合在最终的全连接层中。这两个互补的数据源允许在每个传感器模态内解决模糊性,从而提高了先前方法的准确性。对流行的Human 3.6M数据集(Ionescu等人,Intell IEEE Trans Pattern Anal Mach 36(7):1325-1339,2014)、新发布的TotalCapture数据集和一组具有挑战性的户外视频TotalCaptureOutdoor进行了广泛的评估。我们发布了新的混合MVV数据集(TotalCapture),包括多视点视频,IMU和来自商业运动捕捉系统的精确3D骨骼关节地面实况。该数据集可在http://cvssp.org/data/totalcapture/在线获得。
We propose an approach to accurately estimate 3D human pose by fusing multi-viewpoint video (MVV) with inertial measurement unit (IMU) sensor data, without optical markers, a complex hardware setup or a full body model. Uniquely we use a multi-channel 3D convolutional neural network to learn a pose embedding from visual occupancy and semantic 2D pose estimates from the MVV in a discretised volumetric probabilistic visual hull. The learnt pose stream is concurrently processed with a forward kinematic solve of the IMU data and a temporal model (LSTM) exploits the rich spatial and temporal long range dependencies among the solved joints, the two streams are then fused in a final fully connected layer. The two complementary data sources allow for ambiguities to be resolved within each sensor modality, yielding improved accuracy over prior methods. Extensive evaluation is performed with state of the art performance reported on the popular Human 3.6M dataset(Ionescu et al. in Intell IEEE Trans Pattern Anal Mach 36(7):1325-1339, 2014), the newly released TotalCapture dataset and a challenging set of outdoor videos TotalCaptureOutdoor. We release the new hybrid MVV dataset (TotalCapture) comprising of multi-viewpoint video, IMU and accurate 3D skeletal joint ground truth derived from a commercial motion capture system. The dataset is available online at http://cvssp.org/data/totalcapture/.