课题基金 / 基金详情

Audio-Visual Speech Enhancement and Speaker Separation

Audio-Visual Speech Enhancement and Speaker Separation
视听语音增强和扬声器分离
批准号:
2243852
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The problem with audio perception is that individual sounds are mixed together with unknown acoustic reverberations, and this makes it impossible to extract them without prior knowledge of the source characteristics. The problem of audio-source separation is a fundamental problem in audio perception. Humans have the ability to understanding speech when it is mixed with other types of sound and noise; by isolating and focusing attention to one voice from a multitude. This research aims to reproduce or model this accomplishment of the brain with computational and algorithmic means. Speech enhancement is a method of increasing speech intelligibility by using algorithms to separate and enhance the original source of the speech from others. Automating the process of speech enhancement has many real-world applications such as increasing the effectiveness of assistive technology for the hearing impaired, creating virtual reality with high clarity and better transcription of speech in noisy audio tracks. Additionally, with ever-increasing use of audio-visual and voice-controlled technologies, the ability to capture and enhance a speaker's voice is becoming imperative in the robustness of automatic speech recognition (ASR) systems. These systems tend to infer speech well in quiet environments, but they struggle when background noise is present.Although recently there has been significant advancement in speech separation using deep learning methods, it is still considered a difficult problem due to time-variant input signals and high variability of reverberant sound fields. Traditionally the task of speech enhancement is either performed on audio-only tracks or the combination of audio and video inputs. Deep learning techniques have been applied to challenging tasks such as removing background noise from speech, separating a speaker from multiple speech signals, or more generally separating arbitrary classes of sound from each other. This work will address the shortcomings of the current methods and will explore conditioning speech separation tasks by conditioning on complementary information, such as visual cues from the speaker's lip motions.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cvpr52688.2022.01024
发表时间: 2022-06
期刊: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者: [Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman]
通讯作者: Akam Rahimi;Triantafyllos Afouras;Andrew Zisserman
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
  • 批准号:
    60373031
  • 项目类别:
    面上项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2003
  • 负责人:
    陈越
  • 依托单位: