Learning View-Invariant Sparse Representations for Cross-View Action Recognition

Learning View-Invariant Sparse Representations for Cross-View Action Recognition
复制标题

DOI:
10.1109/iccv.2013.394
复制
发表时间:
2013-12
期刊:
2013 IEEE International Conference on Computer Vision
影响因子:
--
通讯作者:
Jingjing Zheng;Zhuolin Jiang
Jingjing Zheng;Zhuolin Jiang
中科院分区:
其他
文献类型:
--
作者:
Jingjing Zheng;Zhuolin Jiang

文献摘要

被引文献

相似文献

我们提出了一种方法来共同学习一组特定于视图的字典和一个共同的字典跨视图动作识别。特定于视图的字典集是针对特定视图学习的,而公共字典在不同视图之间共享。我们的方法表示视频在每个视图中使用相应的视图特定的字典和共同的字典。更重要的是,它鼓励从同一动作的不同视图拍摄的视频集具有相似的稀疏表示。通过这种方式,我们可以在由特定于视图的字典集跨越的稀疏特征空间中对齐特定于视图的特征,并在由公共字典跨越的稀疏特征空间中传输视图共享特征。同时,通用字典和视图特定字典集之间的不一致性使我们能够分别利用视图特定特征和视图共享特征中编码的鉴别信息。此外,学习的通用字典不仅能够表示来自未见过视图的动作,而且还使我们的方法在半监督设置中有效,其中不存在对应视频并且目标视图中仅存在少数标签。使用多视图IXMAS数据集的大量实验表明,我们的方法优于许多最近的跨视图动作识别方法。
We present an approach to jointly learn a set of view-specific dictionaries and a common dictionary for cross-view action recognition. The set of view-specific dictionaries is learned for specific views while the common dictionary is shared across different views. Our approach represents videos in each view using both the corresponding view-specific dictionary and the common dictionary. More importantly, it encourages the set of videos taken from different views of the same action to have similar sparse representations. In this way, we can align view-specific features in the sparse feature spaces spanned by the view-specific dictionary set and transfer the view-shared features in the sparse feature space spanned by the common dictionary. Meanwhile, the incoherence between the common dictionary and the view-specific dictionary set enables us to exploit the discrimination information encoded in view-specific features and view-shared features separately. In addition, the learned common dictionary not only has the capability to represent actions from unseen views, but also makes our approach effective in a semi-supervised setting where no correspondence videos exist and only a few labels exist in the target view. Extensive experiments using the multi-view IXMAS dataset demonstrate that our approach outperforms many recent approaches for cross-view action recognition.