Spatial-Aware Object Embeddings for Zero-Shot Localization and Classification of Actions

Spatial-Aware Object Embeddings for Zero-Shot Localization and Classification of Actions
复制标题

DOI:
10.1109/iccv.2017.476
复制
发表时间:
2017-07
期刊:
2017 IEEE International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
P. Mettes;Cees G. M. Snoek
P. Mettes;Cees G. M. Snoek
中科院分区:
其他
文献类型:
--
作者:
P. Mettes;Cees G. M. Snoek

文献摘要

被引文献

相似文献

我们的目标是零镜头定位和分类的人的行动在视频中。传统的方法依赖于全局属性或对象分类分数进行零次知识转移,我们的主要贡献是空间感知对象嵌入。为了达到空间感知,我们在免费提供的演员和对象检测器之上构建嵌入。在词嵌入空间中确定对象的相关性,并进一步用估计的空间偏好来执行。除了局部对象感知,我们还将全局对象感知嵌入到我们的嵌入中,以最大限度地提高演员和对象的交互。最后,我们利用对象的位置和大小的空间感知嵌入展示了一个新的时空动作检索方案与复合查询。四个当代动作视频数据集上的动作定位和分类实验支持我们的建议。除了零拍摄定位和分类设置中的最新结果外,我们的空间感知嵌入甚至与最近的监督动作定位替代方案具有竞争力。
We aim for zero-shot localization and classification of human actions in video. Where traditional approaches rely on global attribute or object classification scores for their zero-shot knowledge transfer, our main contribution is a spatial-aware object embedding. To arrive at spatial awareness, we build our embedding on top of freely available actor and object detectors. Relevance of objects is determined in a word embedding space and further enforced with estimated spatial preferences. Besides local object awareness, we also embed global object awareness into our embedding to maximize actor and object interaction. Finally, we exploit the object positions and sizes in the spatial-aware embedding to demonstrate a new spatiotemporal action retrieval scenario with composite queries. Action localization and classification experiments on four contemporary action video datasets support our proposal. Apart from state-of-the-art results in the zero-shot localization and classification settings, our spatial-aware embedding is even competitive with recent supervised action localization alternatives.