Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge

Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge
复制标题

DOI:
10.1109/cvpr42600.2020.00656
复制
发表时间:
2020-04
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Long Zhao;Xi Peng;Yuxiao Chen;M. Kapadia;Dimitris N. Metaxas
Long Zhao;Xi Peng;Yuxiao Chen;M. Kapadia;Dimitris N. Metaxas
中科院分区:
其他
文献类型:
--
作者:
Long Zhao;Xi Peng;Yuxiao Chen;M. Kapadia;Dimitris N. Metaxas

文献摘要

被引文献

相似文献

跨模态知识蒸馏处理的是将知识从一个由较优模态训练的模型(教师)转移到另一个由较弱模态训练的模型(学生)。现有的方法需要配对训练,两种模式下都有例子。然而,从高级模式访问数据可能并不总是可行的。例如,在3D手部姿势估计的情况下,深度图、点云或立体图像通常比RGB图像捕获更好的手部结构,但大多数图像的收集成本很高。在本文中,我们提出了一种新的方案,在教师不可用的目标数据集中训练学生。我们的关键思想是通过将知识建模为学生参数的先验,将从源数据集(包含来自两种模态的成对示例)中学习到的提取出来的跨模态知识推广到目标数据集。我们将我们的方法命名为“跨模态知识泛化”,并证明我们的方案在标准基准数据集上的3D手部姿态估计具有竞争力。
Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superior modalities may not always be feasible. For example, in the case of 3D hand pose estimation, depth maps, point clouds, or stereo images usually capture better hand structures than RGB images, but most of them are expensive to be collected. In this paper, we propose a novel scheme to train the Student in a Target dataset where the Teacher is unavailable. Our key idea is to generalize the distilled cross-modal knowledge learned from a Source dataset, which contains paired examples from both modalities, to the Target dataset by modeling knowledge as priors on parameters of the Student. We name our method "Cross-Modal Knowledge Generalization" and demonstrate that our scheme results in competitive performance for 3D hand pose estimation on standard benchmark datasets.