Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
复制标题

DOI:
10.1109/cvpr52688.2022.01024
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman
Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman

文献摘要

相似文献

本文的目标是使用不同模式的组合在多说话者和噪声环境中进行语音分离和增强。以前的作品在调节时间或静态视觉证据(例如同步嘴唇运动或面部识别)时表现出了良好的性能。在本文中,我们提出了一个基于同步或异步线索的多模态语音分离和增强的统一框架。为此,我们做出以下贡献:(i)我们设计了一种基于 Transformer 的现代架构,旨在融合不同的模态来解决原始波形域中的语音分离任务; (ii) 我们建议单独或结合视觉信息来调节句子的文本内容; (iii) 我们证明了我们的模型对视听同步偏移的鲁棒性; (iv) 我们在完善的基准数据集 LRS2 和 LRS3 上获得了最先进的性能。
The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good performance when conditioning on temporal or static visual evidence such as synchronised lip movements or face identity. In this paper, we present a unified framework for multi-modal speech separation and enhancement based on synchronous or asynchronous cues. To that end we make the following contributions: (i) we design a modern Transformer-based architecture tailored to fuse different modalities to solve the speech separation task in the raw waveform domain; (ii) we propose conditioning on the textual content of a sentence alone or in combination with visual information; (iii) we demonstrate the robustness of our model to audio-visual synchronisation offsets; and, (iv) we obtain state-of-the-art performance on the well-established benchmark datasets LRS2 and LRS3.