HumanEva: Synchronized Video and Motion Capture Dataset and Baseline Algorithm for Evaluation of Articulated Human Motion

HumanEva: Synchronized Video and Motion Capture Dataset and Baseline Algorithm for Evaluation of Articulated Human Motion
复制标题

DOI:
10.1007/s11263-009-0273-6
复制
发表时间:
2010-03-01
影响因子:
19.5
通讯作者:
Black, Michael J.
Black, Michael J.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sigal, Leonid;Balan, Alexandru O.;Black, Michael J.

文献摘要

被引文献

相似文献

虽然在过去的几年里,对人体运动和姿态估计的研究进展迅速,一直没有系统的定量评估竞争的方法,以建立当前的最先进的。我们提出了使用硬件系统,能够捕捉同步视频和地面实况3D运动获得的数据。由此产生的HumanEva数据集包含多个主题,这些主题执行一组预定义的动作,并重复多次。在60 Hz下收集了大约40,000帧的同步运动捕捉和多视图视频(总共产生超过二十五万个图像帧),以及另外37,000个纯运动捕捉数据的时刻。定义了一组标准的误差测量,用于评估2D和3D姿态估计和跟踪算法。我们还描述了一个基线算法的3D铰接跟踪,使用一个相对标准的贝叶斯框架与优化的形式顺序重要性呼吸和退火粒子滤波。在这个基线算法的背景下,我们探讨了各种似然函数,人体运动的先验模型和算法参数的影响。我们的实验表明,图像观察模型和运动先验在性能中起着重要的作用,并且在多视图实验室环境中,在初始化可用的情况下,贝叶斯滤波往往表现良好。数据集和软件提供给研究界。该基础设施将支持新的关节运动和姿态估计算法的开发,将为新方法的评估和比较提供基线,并将有助于建立人类姿态估计和跟踪的当前技术水平。
While research on articulated human motion and pose estimation has progressed rapidly in the last few years, there has been no systematic quantitative evaluation of competing methods to establish the current state of the art. We present data obtained using a hardware system that is able to capture synchronized video and ground-truth 3D motion. The resulting HumanEva datasets contain multiple subjects performing a set of predefined actions with a number of repetitions. On the order of 40,000 frames of synchronized motion capture and multi-view video (resulting in over one quarter million image frames in total) were collected at 60 Hz with an additional 37,000 time instants of pure motion capture data. A standard set of error measures is defined for evaluating both 2D and 3D pose estimation and tracking algorithms. We also describe a baseline algorithm for 3D articulated tracking that uses a relatively standard Bayesian framework with optimization in the form of Sequential Importance Resampling and Annealed Particle Filtering. In the context of this baseline algorithm we explore a variety of likelihood functions, prior models of human motion and the effects of algorithm parameters. Our experiments suggest that image observation models and motion priors play important roles in performance, and that in a multi-view laboratory environment, where initialization is available, Bayesian filtering tends to perform well. The datasets and the software are made available to the research community. This infrastructure will support the development of new articulated motion and pose estimation algorithms, will provide a baseline for the evaluation and comparison of new methods, and will help establish the current state of the art in human pose estimation and tracking.