End-to-End Audiovisual Speech Recognition System With Multitask Learning

End-to-End Audiovisual Speech Recognition System With Multitask Learning
复制标题

DOI:
10.1109/tmm.2020.2975922
复制
发表时间:
2021
影响因子:
7.3
通讯作者:
Fei Tao;C. Busso
Fei Tao;C. Busso
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fei Tao;C. Busso

文献摘要

被引文献

相似文献

自动语音识别(ASR)系统是当前基于语音的系统的关键组件。然而,周围的噪声会严重降低 ASR 系统的性能。解决这个问题的一个有吸引力的解决方案是使用描述嘴唇活动的视觉特征来增强传统的基于音频的 ASR 系统。本文提出了一种新颖的端到端、多任务学习(MTL)、视听 ASR(AV-ASR)系统。该方法的一个关键新颖之处在于 MTL 的使用,其中主要任务是 AV-ASR,次要任务是视听语音活动检测 (AV-VAD)。我们获得了一个强大而准确的视听系统,可以概括各种条件。通过检测具有语音活动的片段,AV-ASR 的性能得到改善,因为其连接主义时间分类 (CTC) 损失函数可以利用 AV-VAD 对齐信息。此外,端到端系统从原始视听输入中学习两个语音任务的有区别的高级表示,从而提供了直接从数据中挖掘信息的灵活性。所提出的架构考虑了模态内部和跨模态的时间动态,提供了一个有吸引力且实用的融合方案。我们在包含不同通道和环境条件的大型视听语料库(超过 60 小时)上评估所提出的方法,并将结果与​​竞争性单任务学习 (STL) 和 MTL 基线进行比较。尽管我们的主要目标是提高 ASR 任务的性能,但实验结果表明,所提出的方法可以在两个语音任务的所有条件下实现最佳性能。除了 AV-ASR 中最先进的性能之外,所提出的解决方案还可以提供有关语音活动的有价值的信息,解决基于语音的应用程序中两个最重要的任务。
An automatic speech recognition (ASR) system is a key component in current speech-based systems. However, the surrounding acoustic noise can severely degrade the performance of an ASR system. An appealing solution to address this problem is to augment conventional audio-based ASR systems with visual features describing lip activity. This paper proposes a novel end-to-end, multitask learning (MTL), audiovisual ASR (AV-ASR) system. A key novelty of the approach is the use of MTL, where the primary task is AV-ASR, and the secondary task is audiovisual voice activity detection (AV-VAD). We obtain a robust and accurate audiovisual system that generalizes across conditions. By detecting segments with speech activity, the AV-ASR performance improves as its connectionist temporal classification (CTC) loss function can leverage from the AV-VAD alignment information. Furthermore, the end-to-end system learns from the raw audiovisual inputs a discriminative high-level representation for both speech tasks, providing the flexibility to mine information directly from the data. The proposed architecture considers the temporal dynamics within and across modalities, providing an appealing and practical fusion scheme. We evaluate the proposed approach on a large audiovisual corpus (over 60 hours), which contains different channel and environmental conditions, comparing the results with competitive single task learning (STL) and MTL baselines. Although our main goal is to improve the performance of our ASR task, the experimental results show that the proposed approach can achieve the best performance across all conditions for both speech tasks. In addition to state-of-the-art performance in AV-ASR, the proposed solution can also provide valuable information about speech activity, solving two of the most important tasks in speech-based applications.