Automatic and Efficient Human Pose Estimation for Sign Language Videos

Automatic and Efficient Human Pose Estimation for Sign Language Videos
复制标题

DOI:
10.1007/s11263-013-0672-6
复制
发表时间:
2014-10-01
影响因子:
19.5
通讯作者:
Zisserman, Andrew
Zisserman, Andrew
中科院分区:
计算机科学2区
文献类型:
--
作者:
Charles, James;Pfister, Tomas;Zisserman, Andrew

文献摘要

被引文献

相似文献

我们提出了一种全自动手臂和手部追踪器,可以在超过一个小时的连续手语视频序列中检测关节位置。为了实现这一目标,我们在四个方面做出了贡献:(i)我们证明了覆盖的签名者可以使用分层模型在所有帧上使用共同分割从背景电视广播中分离出来;(ii)我们展示了关节位置(肩膀,肘部,手腕)可以使用随机森林回归器预测每帧,只给出这个分割和颜色模型;(iii)我们证明随机森林可以从现有的半自动跟踪器训练,但计算成本很高;(iv)引入评估器来评估每帧预测的关节位置是否正确。该方法应用于20个不同背景、具有挑战性的成像条件和不同签名者的签名视频。我们的框架优于Buehler等人的最先进的长期跟踪器(International Journal of Computer Vision 95:180- 197,2011),不需要手动注释该工作,并且在自动初始化之后,实时执行跟踪。与使用Yang和Ramanan的姿态估计方法获得的结果相比,我们也获得了更好的关节定位结果(IEEE计算机视觉和模式识别会议论文集,2011)。
We present a fully automatic arm and hand tracker that detects joint positions over continuous sign language video sequences of more than an hour in length. To achieve this, we make contributions in four areas: (i) we show that the overlaid signer can be separated from the background TV broadcast using co-segmentation over all frames with a layered model; (ii) we show that joint positions (shoulders, elbows, wrists) can be predicted per-frame using a random forest regressor given only this segmentation and a colour model; (iii) we show that the random forest can be trained from an existing semi-automatic, but computationally expensive, tracker; and, (iv) introduce an evaluator to assess whether the predicted joint positions are correct for each frame. The method is applied to 20 signing footage videos with changing background, challenging imaging conditions, and for different signers. Our framework outperforms the state-of-the-art long term tracker by Buehler et al. (International Journal of Computer Vision 95:180-197, 2011), does not require the manual annotation of that work, and, after automatic initialisation, performs tracking in real-time. We also achieve superior joint localisation results to those obtained using the pose estimation method of Yang and Ramanan (Proceedings of the IEEE conference on computer vision and pattern recognition, 2011).