My lips are concealed: Audio-visual speech enhancement through obstructions

My lips are concealed: Audio-visual speech enhancement through obstructions
复制标题

DOI:
10.21437/interspeech.2019-3114
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Triantafyllos Afouras;Joon Son Chung;Andrew Zisserman
Triantafyllos Afouras;Joon Son Chung;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Triantafyllos Afouras;Joon Son Chung;Andrew Zisserman

文献摘要

被引文献

相似文献

我们的目标是一种视听模型,用于将单个扬声器与其他扬声器和背景噪声等混合声音分开。此外,我们希望即使在由于遮挡而暂时不存在视觉提示的情况下也能听到说话者的声音。为此,我们引入了一个深度视听语音增强网络,它能够通过对说话人的嘴唇动作和/或他们的声音表示进行条件处理来分离说话人的声音。语音表示可以通过(I)登记或(Ii)自我登记--在给定足够的无障碍的视觉输入的情况下即时学习表示--来获得。该模型通过混合音频和在嘴部周围引入人工遮挡来训练,以防止视觉通道占据主导地位。该方法是独立于说话人的,我们在训练过程中没有听到(和看不到)的说话人的真实例子中展示了它。该方法还改进了以前的模型,特别是对于视觉通道中的遮挡情况。
Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily absent due to occlusion. To this end we introduce a deep audio-visual speech enhancement network that is able to separate a speaker's voice by conditioning on both the speaker's lip movements and/or a representation of their voice. The voice representation can be obtained by either (i) enrollment, or (ii) by self-enrollment -- learning the representation on-the-fly given sufficient unobstructed visual input. The model is trained by blending audios, and by introducing artificial occlusions around the mouth region that prevent the visual modality from dominating. The method is speaker-independent, and we demonstrate it on real examples of speakers unheard (and unseen) during training. The method also improves over previous models in particular for cases of occlusion in the visual modality.