Recognizing action at a distance

Recognizing action at a distance
复制标题

DOI:
10.1109/iccv.2003.1238420
复制
发表时间:
2003-10
期刊:
Proceedings Ninth IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
Alexei A. Efros;A. Berg;Greg Mori;Jitendra Malik
Alexei A. Efros;A. Berg;Greg Mori;Jitendra Malik
中科院分区:
其他
文献类型:
--
作者:
Alexei A. Efros;A. Berg;Greg Mori;Jitendra Malik

文献摘要

被引文献

相似文献

我们的目标是在远处识别人类行为,其分辨率可能是整个人的高度,例如 30 像素。我们引入了一种新颖的运动描述符,该描述符基于每个稳定人物的时空体积中的光流测量,以及在最近邻框架中使用的相关相似性度量。利用噪声光流测量是关键的挑战,解决这个问题的方法是,不要将光流视为精确的像素位移,而是将其视为噪声测量的空间模式,这些测量经过仔细平滑和聚合以形成我们的时空运动描述符。为了对查询序列中人物执行的动作进行分类,我们从存储的带注释的视频序列数据库中检索最近邻。我们还可以使用这些检索到的样本将 2D/3D 骨架转移到查询序列中的图形上,以及两种形式的基于数据的动作合成“照我做”和“照我说的做”。结果在芭蕾舞、网球和足球数据集上进行了演示。
Our goal is to recognize human action at a distance, at resolutions where a whole person may be, say, 30 pixels tall. We introduce a novel motion descriptor based on optical flow measurements in a spatiotemporal volume for each stabilized human figure, and an associated similarity measure to be used in a nearest-neighbor framework. Making use of noisy optical flow measurements is the key challenge, which is addressed by treating optical flow not as precise pixel displacements, but rather as a spatial pattern of noisy measurements which are carefully smoothed and aggregated to form our spatiotemporal motion descriptor. To classify the action being performed by a human figure in a query sequence, we retrieve nearest neighbor(s) from a database of stored, annotated video sequences. We can also use these retrieved exemplars to transfer 2D/3D skeletons onto the figures in the query sequence, as well as two forms of data-based action synthesis "do as I do" and "do as I say". Results are demonstrated on ballet, tennis as well as football datasets.