Few-shot activity recognition with cross-modal memory network

Few-shot activity recognition with cross-modal memory network
复制标题

DOI:
10.1016/j.patcog.2020.107348
复制
发表时间:
2020-12
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Lingling Zhang;Xiaojun Chang;Jun Liu;Minnan Luo;M. Prakash;A. Hauptmann
Lingling Zhang;Xiaojun Chang;Jun Liu;Minnan Luo;M. Prakash;A. Hauptmann
中科院分区:
其他
文献类型:
--
作者:
Lingling Zhang;Xiaojun Chang;Jun Liu;Minnan Luo;M. Prakash;A. Hauptmann

文献摘要

被引文献

相似文献

基于深度学习的动作识别方法需要大量的标记训练数据。然而,对大规模视频数据进行标记既耗时又繁琐。在本文中,我们考虑了一个更具挑战性的训练样本很少的少镜头动作识别问题。为了解决这个问题,记忆网络被设计成使用外部存储器来记住在训练中学习到的经验,然后在测试过程中将其应用于少量的预测。然而,现有的基于记忆的方法只是在记忆中使用固定的标签嵌入来更新视觉信息,不能很好地适应测试过程中的新活动。为了解决这个问题,我们提出了一种新颖的端到端跨模态记忆网络,用于少镜头活动识别。具体来说,所提出的内存体系结构存储了与人类活动相关的一些高级属性的动态视觉和文本语义。在测试阶段,学习记忆可以为新活动的识别提供有效的多模态信息。在HMDB51和UCF101两个视频数据集上的大量实验结果表明,我们的方法比其他方法有明显的改进。
Deep learning based action recognition methods require large amount of labelled training data. However, labelling large-scale video data is time consuming and tedious. In this paper, we consider a more challenging few-shot action recognition problem where the training samples are few and rare. To solve this problem, memory network has been designed to use an external memory to remember the experience learned in training and then apply it to few-shot prediction during testing. However, existing memory-based methods just update the visual information with fixed label embeddings in the memory, which cannot adapt well to novel activities during testing. To alleviate the issue, we propose a novel end-to-end cross-modal memory network for few-shot activity recognition. Specifically, the proposed memory architecture stores the dynamic visual and textual semantics for some high-level attributes related to human activities. And the learned memory can provide effective multi-modal information for new activity recognition in the testing stage. Extensive experimental results on two video datasets, including HMDB51 and UCF101, indicate that our method could achieve significant improvements over other previous methods.