Semantics-Aware Adaptive Knowledge Distillation for Sensor-to-Vision Action Recognition

Semantics-Aware Adaptive Knowledge Distillation for Sensor-to-Vision Action Recognition
复制标题

DOI:
10.1109/tip.2021.3086590
复制
发表时间:
2021-01-01
影响因子:
10.6
通讯作者:
Lin, Liang
Lin, Liang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Yang;Wang, Keze;Lin, Liang

文献摘要

被引文献

相似文献

现有的基于视觉的动作识别容易受到遮挡和外观变化的影响,而可穿戴传感器可以通过用一维时间序列信号(例如加速度,陀螺仪和方向)捕获人体运动来缓解这些挑战。对于同一个动作,从视觉传感器(视频或图像)和可穿戴传感器学习的知识可能是相关和互补的。然而,可穿戴传感器和视觉传感器捕获的动作数据在数据维度、数据分布和固有信息内容上存在显著的模态差异。在本文中,我们提出了一种新的框架,名为语义感知自适应知识蒸馏网络(SAKDN),以提高视觉传感器模态(视频)的动作识别自适应传输和提取的知识,从多个可穿戴传感器。SAKDN使用多个可穿戴传感器作为教师模式,并使用RGB视频作为学生模式。为了保持局部时间关系并便于使用视觉深度学习模型,我们通过设计基于Gramian角场的虚拟图像生成模型将可穿戴传感器的一维时间序列信号转换为二维图像。然后,我们引入了一种新的保持相似性的自适应多模态融合模块(SPAMFM),自适应地融合来自不同教师网络的中间表示知识。最后,为了充分利用多个训练有素的教师网络的知识并将其转移到学生网络,我们提出了一种新的图形引导的语义判别映射(GSDM)模块,该模块利用图形引导的消融分析来产生良好的视觉解释,以突出显示跨模态的重要区域,并同时保留原始数据的相互关系。Berkeley-MHAD,UTD-MHAD和MMAct数据集上的实验结果很好地证明了我们提出的SAKDN从可穿戴传感器模态到视觉传感器模态的自适应知识转移的有效性。该代码可在https://github.com/YangLiu9208/SAKDN上公开获取。
Existing vision-based action recognition is susceptible to occlusion and appearance variations, while wearable sensors can alleviate these challenges by capturing human motion with one-dimensional time-series signals (e.g. acceleration, gyroscope, and orientation). For the same action, the knowledge learned from vision sensors (videos or images) and wearable sensors, may be related and complementary. However, there exists a significantly large modality difference between action data captured by wearable-sensor and vision-sensor in data dimension, data distribution, and inherent information content. In this paper, we propose a novel framework, named Semantics-aware Adaptive Knowledge Distillation Networks (SAKDN), to enhance action recognition in vision-sensor modality (videos) by adaptively transferring and distilling the knowledge from multiple wearable sensors. The SAKDN uses multiple wearable-sensors as teacher modalities and uses RGB videos as student modalities. To preserve the local temporal relationship and facilitate employing visual deep learning models, we transform one-dimensional time-series signals of wearable sensors to two-dimensional images by designing a gramian angular field based virtual image generation model. Then, we introduce a novel Similarity-Preserving Adaptive Multi-modal Fusion Module (SPAMFM) to adaptively fuse intermediate representation knowledge from different teacher networks. Finally, to fully exploit and transfer the knowledge of multiple well-trained teacher networks to the student network, we propose a novel Graph-guided Semantically Discriminative Mapping (GSDM) module, which utilizes graph-guided ablation analysis to produce a good visual explanation to highlight the important regions across modalities and concurrently preserve the interrelations of original data. Experimental results on Berkeley-MHAD, UTD-MHAD, and MMAct datasets well demonstrate the effectiveness of our proposed SAKDN for adaptive knowledge transfer from wearable-sensors modalities to vision-sensors modalities. The code is publicly available at https://github.com/YangLiu9208/SAKDN.