Multi-Modal Domain Adaptation for Fine-Grained Action Recognition

Multi-Modal Domain Adaptation for Fine-Grained Action Recognition
复制标题

DOI:
10.1109/cvpr42600.2020.00020
复制
发表时间:
2020-01
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jonathan Munro;D. Damen
Jonathan Munro;D. Damen
中科院分区:
其他
文献类型:
--
作者:
Jonathan Munro;D. Damen

文献摘要

被引文献

相似文献

细粒度动作识别数据集表现出环境偏差,其中多个视频序列是从有限数量的环境中捕获的。在一个环境中训练模型并在另一个环境中部署,由于不可避免的域转移,导致性能下降。无监督域自适应(UDA)方法经常使用源域和目标域之间的对抗训练。然而,这些方法没有探索每个域中视频的多模态性质。在这项工作中,我们利用模态的对应性作为UDA的自监督对齐方法,以及对抗对齐(图1)。我们使用两种常用的动作识别模式:RGB和光流,在来自大规模EPIC-Kitterfly数据集的三个厨房上测试我们的方法。我们发现,多模态自我监督比仅源培训平均提高了2.4%的性能。然后,我们将联合收割机对抗训练与多模态自我监督相结合,表明我们的方法比其他UDA方法高出3%。
Fine-grained action recognition datasets exhibit environmental bias, where multiple video sequences are captured from a limited number of environments. Training a model in one environment and deploying in another results in a drop in performance due to an unavoidable domain shift. Unsupervised Domain Adaptation (UDA) approaches have frequently utilised adversarial training between the source and target domains. However, these approaches have not explored the multi-modal nature of video within each domain. In this work we exploit the correspondence of modalities as a self-supervised alignment approach for UDA in addition to adversarial alignment (Fig. 1). We test our approach on three kitchens from the large-scale EPIC-Kitchens dataset, using two modalities commonly employed for action recognition: RGB and Optical Flow. We show that multi-modal self-supervision alone improves the performance over source-only training by 2.4% on average. We then combine adversarial training with multi-modal self-supervision, showing that our approach outperforms other UDA methods by 3%.