Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition

Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition
复制标题

DOI:
10.1109/cvpr52688.2022.01431
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Lijin Yang;Yifei Huang;Yusuke Sugano;Y. Sato
Lijin Yang;Yifei Huang;Yusuke Sugano;Y. Sato
中科院分区:
其他
文献类型:
--
作者:
Lijin Yang;Yifei Huang;Yusuke Sugano;Y. Sato

文献摘要

相似文献

无监督域自适应视频动作识别的目的是使用仅用域外(源)注释训练的模型来识别目标域的动作。视频的固有复杂性使得这项任务具有挑战性,但也为利用多模式输入(例如,RGB、流、音频)。大多数以前的作品利用多模态信息,要么单独对齐每个模态,要么通过跨模态自我监督学习表示。与以前的工作不同,我们发现跨域对齐可以更有效地通过使用跨模态相互作用第一。跨模态知识交互允许其他模态补充缺失的可转移信息,因为跨模态的互补性。此外,可以使用跨模态共识来突出数据的最可转移方面。在这项工作中,我们提出了一种新的模型,共同考虑这两个领域的自适应动作识别的特点。我们实现这一点,通过实施两个模块,其中第一个模块交换互补的可转移信息跨模态通过语义空间,和第二个模块找到最可转移的空间区域的共识的基础上,所有模态。大量的实验验证了我们提出的方法在多个基准数据集上的性能明显优于最先进的方法,包括复杂的细粒度数据集EPIC-Kitchen-100。
Unsupervised domain adaptive video action recognition aims to recognize actions of a target domain using a model trained with only out-of-domain (source) annotations. The inherent complexity of videos makes this task challenging but also provides ground for leveraging multi-modal inputs (e.g., RGB, Flow, Audio). Most previous works utilize the multi-modal information by either aligning each modality individually or learning representation via cross-modal self-supervision. Different from previous works, we find that the cross-domain alignment can be more effectively done by using cross-modal interaction first. Cross-modal knowledge interaction allows other modalities to supplement missing transferable information because of the cross-modal complementarity. Also, the most transferable aspects of data can be highlighted using cross-modal consensus. In this work, we present a novel model that jointly considers these two characteristics for domain adaptive action recognition. We achieve this by implementing two modules, where the first module exchanges complementary transferable information across modalities through the semantic space, and the second module finds the most transferable spatial region based on the consensus of all modalities. Extensive experiments validate that our proposed method can significantly outperform state-of-the-art methods on multiple benchmark datasets, including the complex fine-grained dataset EPIC-Kitchens-100.