课题基金 / 基金详情

Audio-Visual Speech Enhancement and Speaker Separation

Audio-Visual Speech Enhancement and Speaker Separation
视听语音增强和扬声器分离
批准号:
2243852
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
音频感知的问题在于,单个声音与未知的混响混合在一起,这使得在没有事先了解源特性的情况下无法提取它们。音源分离问题是音频感知中的一个基本问题。当语音与其他类型的声音和噪音混合在一起时,人类有能力理解语音;把注意力从众多声音中分离出来,集中到一个声音上。这项研究旨在用计算和算法的手段再现或模拟大脑的这一成就。语音增强是一种提高语音清晰度的方法,通过使用算法将原始语音源与其他语音源分离和增强。语音增强过程的自动化在现实世界中有许多应用,例如提高听力障碍患者辅助技术的有效性,创造高清晰度的虚拟现实,以及在嘈杂的音轨中更好地转录语音。此外,随着视听和语音控制技术的不断增加,捕捉和增强说话者声音的能力在自动语音识别(ASR)系统的稳健性中变得至关重要。这些系统倾向于在安静的环境中很好地推断语音,但当背景噪音存在时,它们就会遇到困难。尽管近年来在使用深度学习方法进行语音分离方面取得了重大进展,但由于输入信号时变和混响声场的高可变性,语音分离仍然被认为是一个难题。传统的语音增强任务要么在音频轨道上进行,要么在音频和视频输入的组合上进行。深度学习技术已被应用于具有挑战性的任务,例如从语音中去除背景噪声,从多个语音信号中分离说话者,或者更广泛地将任意类别的声音彼此分离。这项工作将解决当前方法的缺点,并将探索通过对互补信息(如说话者嘴唇运动的视觉线索)进行条件反射来调节语音分离任务。
英文摘要
The problem with audio perception is that individual sounds are mixed together with unknown acoustic reverberations, and this makes it impossible to extract them without prior knowledge of the source characteristics. The problem of audio-source separation is a fundamental problem in audio perception. Humans have the ability to understanding speech when it is mixed with other types of sound and noise; by isolating and focusing attention to one voice from a multitude. This research aims to reproduce or model this accomplishment of the brain with computational and algorithmic means. Speech enhancement is a method of increasing speech intelligibility by using algorithms to separate and enhance the original source of the speech from others. Automating the process of speech enhancement has many real-world applications such as increasing the effectiveness of assistive technology for the hearing impaired, creating virtual reality with high clarity and better transcription of speech in noisy audio tracks. Additionally, with ever-increasing use of audio-visual and voice-controlled technologies, the ability to capture and enhance a speaker's voice is becoming imperative in the robustness of automatic speech recognition (ASR) systems. These systems tend to infer speech well in quiet environments, but they struggle when background noise is present.Although recently there has been significant advancement in speech separation using deep learning methods, it is still considered a difficult problem due to time-variant input signals and high variability of reverberant sound fields. Traditionally the task of speech enhancement is either performed on audio-only tracks or the combination of audio and video inputs. Deep learning techniques have been applied to challenging tasks such as removing background noise from speech, separating a speaker from multiple speech signals, or more generally separating arbitrary classes of sound from each other. This work will address the shortcomings of the current methods and will explore conditioning speech separation tasks by conditioning on complementary information, such as visual cues from the speaker's lip motions.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr52688.2022.01024
发表时间: 2022-06
期刊: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者: [Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman]
通讯作者: Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
  • 批准号:
    60373031
  • 项目类别:
    面上项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2003
  • 负责人:
    陈越
  • 依托单位: