Learning from the Master: Distilling Cross-modal Advanced Knowledge for Lip Reading

Learning from the Master: Distilling Cross-modal Advanced Knowledge for Lip Reading
复制标题

DOI:
10.1109/cvpr46437.2021.01312
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Sucheng Ren;Yong Du;Jian Lv;Guoqiang Han;Shengfeng He
Sucheng Ren;Yong Du;Jian Lv;Guoqiang Han;Shengfeng He
中科院分区:
其他
文献类型:
--
作者:
Sucheng Ren;Yong Du;Jian Lv;Guoqiang Han;Shengfeng He

文献摘要

被引文献

相似文献

唇读阅读的目标是从无声的唇视频中预测出口语句子。由于这样的视觉任务通常比其对应的语音识别执行得更差,一个潜在的方案是从通过音频信号预训练的教师中提取知识。然而,跨模态数据之间的潜在域间隙可能导致学习歧义,从而限制唇阅读的性能。本文提出了一个新的唇阅读合作框架,并考虑了两个方面的问题:1)教师应理解双通道知识,以弥合固有的跨通道鸿沟; 2)教师应根据学生的发展自适应地调整教学内容。为此,我们引入了一个可训练的“主”网络,它摄取音频信号和无声的嘴唇视频,而不是一个预先训练的老师。主设备从三种特征模态产生logit:音频模态、视频模态及其组合。为了进一步提供一个互动的策略,有机地融合这些知识,我们经常与特定任务的反馈,从学生,其中隐含的要求嵌入的主人。同时,我们在系统中加入了几个导师网络作为指导,以灵活地强调丰富的知识。此外,我们纳入了课程学习设计,以确保更好的衔接。大量的实验表明,所提出的网络优于国家的最先进的方法在几个基准,包括在单词级和词典级的情况下。
Lip reading aims to predict the spoken sentences from silent lip videos. Due to the fact that such a vision task usually performs worse than its counterpart speech recognition, one potential scheme is to distill knowledge from a teacher pretrained by audio signals. However, the latent domain gap between the cross-modal data could lead to a learning ambiguity and thus limits the performance of lip reading. In this paper, we propose a novel collaborative framework for lip reading, and two aspects of issues are considered: 1) the teacher should understand bi-modal knowledge to possibly bridge the inherent cross-modal gap; 2) the teacher should adjust teaching contents adaptively with the evolution of the student. To these ends, we introduce a trainable "master" network which ingests both audio signals and silent lip videos instead of a pretrained teacher. The master produces logits from three modalities of features: audio modality, video modality, and their combination. To further provide an interactive strategy to fuse these knowledge organically, we regularize the master with the task-specific feedback from the student, in which the requirement of the student is implicitly embedded. Meanwhile, we involve a couple of "tutor" networks into our system as guidance for emphasizing the fruitful knowledge flexibly. In addition, we incorporate a curriculum learning design to ensure a better convergence. Extensive experiments demonstrate that the proposed network outperforms the state-of-the-art methods on several benchmarks, including in both word-level and sentence-level scenarios.