The Conversation: Deep Audio-Visual Speech Enhancement

The Conversation: Deep Audio-Visual Speech Enhancement
复制标题

DOI:
10.21437/interspeech.2018-1400
复制
发表时间:
2018-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Triantafyllos Afouras;Joon Son Chung;Andrew Zisserman
Triantafyllos Afouras;Joon Son Chung;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Triantafyllos Afouras;Joon Son Chung;Andrew Zisserman

文献摘要

相似文献

我们的目标是从视频中的多人同时语音中分离出单个扬声器。在这一领域的现有作品集中在试图从已知的扬声器在受控环境中分离的话语。在本文中,我们提出了一种深度视听语音增强网络,它能够通过预测目标信号的幅度和相位来分离相应视频中给定嘴唇区域的说话者的声音。该方法适用于在训练过程中听不见和看不见的说话者,以及不受约束的环境。我们展示了强有力的定量和定性结果,隔离了极具挑战性的现实世界的例子。
Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose a deep audio-visual speech enhancement network that is able to separate a speaker's voice given lip regions in the corresponding video, by predicting both the magnitude and the phase of the target signal. The method is applicable to speakers unheard and unseen during training, and for unconstrained environments. We demonstrate strong quantitative and qualitative results, isolating extremely challenging real-world examples.