Signs in time: Encoding human motion as a temporal image

Signs in time: Encoding human motion as a temporal image
复制标题

DOI:
--
复制
发表时间:
2016-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Joon Son Chung;Andrew Zisserman
Joon Son Chung;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Joon Son Chung;Andrew Zisserman

文献摘要

被引文献

相似文献

这项工作的目标是识别和定位图像时间序列中的短时信号,其中强有力的监督是不可用的训练。为此,我们提出了一种图像编码,它以适合使用ConvNet学习的形式简洁地表示视频序列中的人体运动。编码将图像中的姿态信息减少到一列,大大减少了网络的输入要求,但保留了识别的基本信息。该编码应用于识别和本地化英国手语(BSL)视频中的手势。我们证明,使用建议的编码,信号短至10帧的持续时间可以从剪辑持续数百帧,只使用弱(剪辑级)监督和相当大的标签噪声。
The goal of this work is to recognise and localise short temporal signals in image time series, where strong supervision is not available for training. To this end we propose an image encoding that concisely represents human motion in a video sequence in a form that is suitable for learning with a ConvNet. The encoding reduces the pose information from an image to a single column, dramatically diminishing the input requirements for the network, but retaining the essential information for recognition. The encoding is applied to the task of recognizing and localizing signed gestures in British Sign Language (BSL) videos. We demonstrate that using the proposed encoding, signs as short as 10 frames duration can be learnt from clips lasting hundreds of frames using only weak (clip level) supervision and with considerable label noise.