课题基金 / 基金详情

Multi-Modal Blind Source Separation for Robot Audition

Multi-Modal Blind Source Separation for Robot Audition
机器人试镜的多模态盲源分离
批准号:
EP/H012842/1
负责人:
Wenwu Wang
金额:
$14.69万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2009
资助国家:
英国
项目状态:
已结题
起止时间:
2009 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该提案借鉴了萨里大学视觉、语音和信号处理中心在盲源分离和多模式(视听)语音处理方面的专业知识。目标是在室内环境中存在多个竞争声源的情况下执行目标语音的源分离,从而最终提供在不受控制的自然环境中对听觉场景的自动机器感知的进展。这项工作的基本新奇是利用视觉线索,以提高频域盲源分离算法的操作。利用这样的视听处理的目标在于减轻置换问题、欠定问题(即,当源的数量大于麦克风的数量时)和混响问题,这些问题目前限制了盲源分离算法的实用性。因此,工作的重点是可用于执行声音信号自动分离的信号处理算法和软件工具,例如,对于一个机器人。这项提案中的工作以调查人员的丰富经验为基础,其中两人来自盲源分离和数字语音处理领域,另一人来自计算机视觉和模式识别领域。拟议研究的成果将对英国国防工业具有相当大的价值,特别是在目标分离,检测和多路径缓解(或去混响)领域,例如,人机交互,安全监视和人机交互等应用。
英文摘要
This proposal draws on expertise in blind source separation and multimodal (audio-visual) speech processing within the Centre for Vision Speech and Signal Processing at University of Surrey. The objective is to perform source separation of the target speech in the presence of multiple competing sound sources in room environments and thereby ultimately provide progress towards automatic machine perception of auditory scenes within an un-controlled natural environment. The fundamental novelty in this work is to exploit visual cues for enhancing the operation of frequency domain blind source separation algorithms. Exploitation of such audio-visual processing is targeted at mitigating the permutation problem, the underdetermined problem (i.e. when the number of sources is greater than the number of microphones), and the reverberation problem, which currently limits the practical applicability of blind source separation algorithms. The focus of the work is therefore on the signal processing algorithms and software tools that can be used to perform automatic separation of sound signals, e.g., for a robot. The body of work in this proposal is underpinned by the substantial experience of the investigators, two from the areas of blind source separation and digital speech processing, and one from the area of computer vision and pattern recognition. The outcomes of the proposed research will be of considerable value to the UK defence industry working especially in the areas of target separation, detection and multi-path mitigation (or dereverberation), with applications in, for example, human-robot interaction, security surveillance and human-computer interaction.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.21437/interspeech.2010-192
发表时间: 2010-09
期刊:
影响因子: --
作者: [Qingju Liu;Wenwu Wang;P. Jackson]
通讯作者: Qingju Liu;Wenwu Wang;P. Jackson
DOI: 10.1109/taslp.2014.2320637
发表时间: 2014-09
期刊: IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子: --
作者: [Atiyeh Alinaghi;P. Jackson;Qingju Liu;Wenwu Wang]
通讯作者: Atiyeh Alinaghi;P. Jackson;Qingju Liu;Wenwu Wang
Interference Reduction in Reverberant <newline/>Speech Separation With Visual <newline/>Voice Activity Detection
通过视觉 <newline/> 语音活动检测来减少混响 <newline/> 语音分离中的干扰
DOI: 10.1109/tmm.2014.2322824
发表时间: 2014
期刊: IEEE Transactions on Multimedia
影响因子: 7.3
作者: [Liu Q]
通讯作者: Liu Q
DOI: 10.1109/tsp.2013.2277834
发表时间: 2013-11
期刊: IEEE Transactions on Signal Processing
影响因子: 5.4
作者: [Qingju Liu;Wenwu Wang;P. Jackson;M. Barnard;J. Kittler;J. Chambers]
通讯作者: Qingju Liu;Wenwu Wang;P. Jackson;M. Barnard;J. Kittler;J. Chambers
海外基金