Key Frame Proposal Network for Efficient Pose Estimation in Videos

Key Frame Proposal Network for Efficient Pose Estimation in Videos
复制标题

DOI:
10.1007/978-3-030-58520-4_36
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuexi Zhang;Yin Wang;O. Camps;M. Sznaier
Yuexi Zhang;Yin Wang;O. Camps;M. Sznaier
中科院分区:
其他
文献类型:
--
作者:
Yuexi Zhang;Yin Wang;O. Camps;M. Sznaier

文献摘要

相似文献

视频中的人体姿态估计依赖于本地信息,要么独立地估计每个帧,要么跨帧跟踪姿态。在本文中,我们提出了一种新的方法相结合的局部方法与全球背景。我们引入了一个轻量级的,无监督的,关键帧建议网络(K-FPN)来选择信息帧和一个学习字典来恢复这些帧的整个姿势序列。K-FPN加速了姿态估计,并对具有遮挡、运动模糊和照明变化的坏帧提供了鲁棒性,而学习的字典提供了全局动态上下文。在Penn Action和sub-JHMDB数据集上的实验表明,该方法具有较高的准确率和较好的加速性能。
Human pose estimation in video relies on local information by either estimating each frame independently or tracking poses across frames. In this paper, we propose a novel method combining local approaches with global context. We introduce a light weighted, unsupervised, key frame proposal network (K-FPN) to select informative frames and a learned dictionary to recover the entire pose sequence from these frames. The K-FPN speeds up the pose estimation and provides robustness to bad frames with occlusion, motion blur, and illumination changes, while the learned dictionary provides global dynamic context. Experiments on Penn Action and sub-JHMDB datasets show that the proposed method achieves state-of-the-art accuracy, with substantial speed-up.