Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments

Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments
复制标题

DOI:
10.1109/tpami.2013.248
复制
发表时间:
2014-07-01
影响因子:
23.6
通讯作者:
Sminchisescu, Cristian
Sminchisescu, Cristian
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ionescu, Catalin;Papava, Dragos;Sminchisescu, Cristian

文献摘要

被引文献

相似文献

我们引入了一个新数据集 Human3.6M,其中包含 360 万个准确的 3D 人体姿势,这些姿势是通过记录 5 名女性和 6 名男性受试者在 4 个不同视角下的表现而获得的,用于训练真实的人体传感系统并评估下一代人体姿势估计模型和算法。除了将当前最先进的数据集的大小增加几个数量级之外,我们还旨在通过典型人类活动(拍照、打电话、摆姿势、打招呼、吃饭等)中遇到的各种动作和姿势来补充这些数据集,附加同步图像、人体动作捕捉和飞行时间(深度)数据,以及对所有相关主体演员的精确 3D 身体扫描。我们还提供受控混合现实评估场景,其中 3D 人体模型使用动作捕捉进行动画处理,并使用正确的 3D 几何形状插入,在复杂的真实环境中,使用移动摄像机查看,并在遮挡下。最后,我们为数据集提供了一套大规模统计模型和详细的评估基线,说明了其多样性以及研究界未来工作的改进范围。我们的实验表明,与该问题的现有最大公共数据集规模的训练集相比,我们最好的大规模模型可以利用我们的完整训练集获得 20% 的性能提升。然而,通过利用我们的大型数据集来利用更高容量、更复杂的模型,改进的潜力要大得多,应该会刺激未来的研究。该数据集以及相关大规模学习模型、功能、可视化工具以及评估服务器的代码可在线获取:http://vision.imar.ro/ human3.6m。
We introduce a new dataset, Human3.6M, of 3.6 Million accurate 3D Human poses, acquired by recording the performance of 5 female and 6 male subjects, under 4 different viewpoints, for training realistic human sensing systems and for evaluating the next generation of human pose estimation models and algorithms. Besides increasing the size of the datasets in the current state-of-the-art by several orders of magnitude, we also aim to complement such datasets with a diverse set of motions and poses encountered as part of typical human activities (taking photos, talking on the phone, posing, greeting, eating, etc.), with additional synchronized image, human motion capture, and time of flight (depth) data, and with accurate 3D body scans of all the subject actors involved. We also provide controlled mixed reality evaluation scenarios where 3D human models are animated using motion capture and inserted using correct 3D geometry, in complex real environments, viewed with moving cameras, and under occlusion. Finally, we provide a set of large-scale statistical models and detailed evaluation baselines for the dataset illustrating its diversity and the scope for improvement by future work in the research community. Our experiments show that our best large-scale model can leverage our full training set to obtain a 20% improvement in performance compared to a training set of the scale of the largest existing public dataset for this problem. Yet the potential for improvement by leveraging higher capacity, more complex models with our large dataset, is substantially vaster and should stimulate future research. The dataset together with code for the associated large-scale learning models, features, visualization tools, as well as the evaluation server, is available online at http://vision.imar.ro/human3.6m.