Lip-Reading Driven Deep Learning Approach for Speech Enhancement

Lip-Reading Driven Deep Learning Approach for Speech Enhancement
复制标题

DOI:
10.1109/tetci.2019.2917039
复制
发表时间:
2021-06-01
影响因子:
5.3
通讯作者:
Whitmer, William M.
Whitmer, William M.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Adeel, Ahsan;Gogate, Mandar;Whitmer, William M.

文献摘要

被引文献

相似文献

本文提出了一种新的唇读驱动的深度学习语音增强框架。与仅依赖于深度学习的基准方法相比,该方法利用了深度学习和分析声学建模(基于滤波的方法)的互补优势。建议的视听语音增强框架在两个层面上运作。在第一级中,采用了一种新的基于深度学习的唇读回归模型。在第二个层次中,唇读近似干净的音频功能被利用,使用增强的,视觉推导的维纳滤波器(EVWF),用于估计干净的音频功率谱。具体地,基于堆叠的长短期记忆(LSTM)的唇读回归模型被设计用于仅使用时间视觉特征(即,唇阅读)。对于干净的语音谱估计,一个新的滤波器组域EVWF制定,它利用估计的语音特征。使用理想AV映射和LSTM驱动的AV映射方法将EVWF与传统的谱减法和对数最小均方误差方法进行比较。建议的AV语音增强框架的潜力进行评估下,四个不同的动态现实世界的情况下[咖啡馆,街道路口,公共交通,和行人区]在不同的SNR水平(从低到高的SNR)使用基准网格和ChiME3语料库。对于客观测试,使用语音质量的感知评估来评估恢复语音的质量。对于主观测试,使用标准的平均意见得分方法与推理统计。比较仿真结果表明显着的唇读和语音增强的语音质量和语音清晰度方面的改善。正在进行的工作旨在增强深度学习驱动的唇读模型的准确性和泛化能力,使用AV线索的上下文整合,从而实现上下文感知的自主AV语音增强。
This paper proposes a novel lip-reading driven deep learning framework for speech enhancement. The approach leverages the complementary strengths of both deep learning and analytical acoustic modeling (filtering-based approach) as compared to benchmark approaches that rely only on deep learning. The proposed audio-visual (AV) speech enhancement framework operates at two levels. In the first level, a novel deep learning based lip-reading regression model is employed. In the second level, lip-reading approximated clean-audio features are exploited, using an enhanced, visually-derived Wiener filter (EVWF), for estimating the clean audio power spectrum. Specifically, a stacked long-short-term memory (LSTM) based lip-reading regression model is designed for estimating the clean audio features using only temporal visual features (i.e., lip reading), by considering a range of prior visual frames. For clean speech spectrum estimation, a new filterbank-domain EVWF is formulated, which exploits the estimated speech features. The EVWF is compared with conventional spectral subtraction and log-minimum mean-square error methods using both ideal AV mapping and LSTM driven AV mapping approaches. The potential of the proposed AV speech enhancement framework is evaluated under four different dynamic real-world scenarios [cafe, street junction, public transport, and pedestrian area] at different SNR levels (ranging from low to high SNRs) using benchmark grid and ChiME3 corpora. For objective testing, perceptual evaluation of speech quality is used to evaluate the quality of restored speech. For subjective testing, the standard mean-opinion-score method is used with inferential statistics. Comparative simulation results demonstrate significant lip-reading and speech enhancement improvements in terms of both speech quality and speech intelligibility. Ongoing work is aimed at enhancing the accuracy and generalization capability of the deep learning driven lip-reading model, using contextual integration of AV cues, leading to context-aware, autonomous AV speech enhancement.