Progressive Cross-modal Knowledge Distillation for Human Action Recognition

Progressive Cross-modal Knowledge Distillation for Human Action Recognition
复制标题

DOI:
10.1145/3503161.3548238
复制
发表时间:
2022-08
期刊:
Proceedings of the 30th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Jianyuan Ni;A. Ngu;Yan Yan-Yan
Jianyuan Ni;A. Ngu;Yan Yan-Yan
中科院分区:
其他
文献类型:
--
作者:
Jianyuan Ni;A. Ngu;Yan Yan-Yan

文献摘要

被引文献

相似文献

基于可穿戴传感器的人类动作识别(HAR)最近取得了显着的成功。然而,基于传感器的可穿戴HAR的准确性性能仍然远远落后于基于视觉模态的系统(即,RGB视频、骨架和深度)。不同的输入方式可以提供互补的线索,从而提高HAR的准确性性能,但如何利用多模态数据的可穿戴传感器的HAR很少被探索。目前,可穿戴设备,即,智能手表只能捕获有限种类的非视觉模态数据。这阻碍了多模态HAR关联,因为它不能同时使用视觉和非视觉模态数据。另一个主要挑战在于如何在计算资源有限的可穿戴设备上有效地利用多模态数据。在这项工作中,我们提出了一种新的渐进式传感器到传感器知识蒸馏(PSKD)模型,它只利用时间序列数据,即,加速度计数据,从智能手表解决可穿戴传感器为基础的HAR问题。具体来说,我们构建多个教师模型使用的数据从教师(人体骨骼序列)和学生(时间序列加速度计数据)的方式。此外,我们提出了一个有效的渐进式学习计划,以消除教师和学生模型之间的性能差距。我们还设计了一个新的损失函数,称为自适应置信语义(ACS),允许学生模型自适应地选择其中一个教师模型或它需要模仿的地面真实标签。为了证明我们提出的PSKD方法的有效性,我们在Berkeley-MHAD,UTD-MHAD和MMAct数据集上进行了广泛的实验。实验结果表明,与以往的基于单传感器的HAR方法相比,所提出的PSKD方法具有较好的性能。
Wearable sensor-based Human Action Recognition (HAR) has achieved remarkable success recently. However, the accuracy performance of wearable sensor-based HAR is still far behind the ones from the visual modalities-based system (i.e., RGB video, skeleton and depth). Diverse input modalities can provide complementary cues and thus improve the accuracy performance of HAR, but how to take advantage of multi-modal data on wearable sensor-based HAR has rarely been explored. Currently, wearable devices, i.e., smartwatches, can only capture limited kinds of non-visual modality data. This hinders the multi-modal HAR association as it is unable to simultaneously use both visual and non-visual modality data. Another major challenge lies in how to efficiently utilize multi-modal data on wearable devices with their limited computation resources. In this work, we propose a novel Progressive Skeleton-to-sensor Knowledge Distillation (PSKD) model which utilizes only time-series data, i.e., accelerometer data, from a smartwatch for solving the wearable sensor-based HAR problem. Specifically, we construct multiple teacher models using data from both teacher (human skeleton sequence) and student (time-series accelerometer data) modalities. In addition, we propose an effective progressive learning scheme to eliminate the performance gap between teacher and student models. We also designed a novel loss function called Adaptive-Confidence Semantic (ACS), to allow the student model to adaptively select either one of the teacher models or the ground-truth label it needs to mimic. To demonstrate the effectiveness of our proposed PSKD method, we conduct extensive experiments on Berkeley-MHAD, UTD-MHAD and MMAct datasets. The results confirm that the proposed PSKD method has competitive performance compared to the previous mono sensor-based HAR methods.