课题基金 / 基金详情

Multi-Modal Blind Source Separation Algorithms

Multi-Modal Blind Source Separation Algorithms
多模态盲源分离算法
批准号:
EP/C535308/2
负责人:
Jonathon Chambers
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --

项目摘要

项目成果

Jonathon Chambers的其他基金

相似基金

相关文献

中文摘要
翻译
该项目涉及模拟人类在办公室环境中将一个语音源与其他扬声器的背景以及可能的噪声源(如空调机组)分离的能力。这就是所谓的鸡尾酒会问题,这是一个非常具有挑战性的任务,使用多个麦克风与计算机一起处理录音,从而提取感兴趣的扬声器。作为人类,我们使用的不仅仅是我们两个耳朵所感知的声音来解决这个问题。例如,我们的眼睛也提供视觉线索,帮助这个过程。因此,这项工作的重点是将从办公室内的麦克风和摄像机获得的音频和视频测量结果结合起来,以帮助分离过程。人类也可能在分离过程中利用语言知识,因此我们计划在分离过程中利用音频和语音记录的数学模型,这些被称为耦合(或融合)隐马尔可夫模型。当在房间内说出一个词时,声波通过许多路径传播到麦克风,这是由于墙壁、天花板或地板或房间内的其他物体(如桌子)上的反射。这种所谓的多径传播是通过所谓的卷积混合来建模的。卷积模型是一个线性的,可能是多通道的系统的输入和输出之间的关系,它记住(有记忆)过去的输入和可能的输出(在这个项目中只有输入)。因此,为了执行分离过程,有必要使用卷积模型。这样的模型将需要许多计算来执行分离,但是这在频域中变得容易得多。我们认为,频域分离是解决这一问题的前进方向,但仍有一些问题有待解决。特别是,如何在时域中重建提取的语音信号(所谓的置换问题),如何处理房间中有两个以上扬声器的情况以及扬声器移动时的情况。我们在这项工作中的方法是使用额外的视觉信息来克服这些问题。因此,我们希望在卡迪夫工程学院的数字信号处理中心内配备一个智能办公室,配备麦克风和摄像机以及必要的计算设施,以记录音频和视频信号的示例,用于测试,最初有两个位置良好的扬声器发出不同的声音,如元音和同意,然后转向更多的扬声器和运动;最终,记录自然的连续语音。总体目标是能够展示在智能办公室内分离任何一个说话者话语的能力,这将促进例如与远程位置处的语音识别器或第三方的交互,如在电话会议中。
英文摘要
This project concerns the emulation of the ability of a human to separate one speech source from a background of other speakers and possibly noise sources, such as an air conditioning unit, within an office environment. This is termed the cocktail party problem and it is a very challenging task to use a number of microphones together with a computer to process the recordings, and thereby extract the speaker of interest. As humans, we use much more than the sound that is perceived by our two ears to address this problem. Our eyes, for example, also provide visual cues which help in the process. It is therefore the focus of this work to integrate both audio and visual measurements, attained from microphones and cameras within the office, to aid in the separation process. The human is also likely to exploit knowledge of language in the separation process, we therefore plan to utilize mathematical models of the audio and speech recordings in the separation process, these are called coupled (or fused) Hidden Markov Models. When a word is uttered within a room, the sound wave propagates through many paths to the microphone, due to reflections on the walls, ceiling or floor, or other objects in the room, such as a table. This so-called multipath propagation, is modelled by what is called a convolutive mixture. A convolutive model is the relationship between the input and output of a linear, possibly multichannel, system which remembers (has memory) past inputs and possibly outputs (only inputs in this project). To perform the separation process it is therefore necessary to use a convolutive model. Such a model would need many calculations to perform separation but this becomes much easier in the frequency domain. Separation in the frequency domain is, we believe, the way forward to tackle this problem, but there are problems to be solved. In particular, how to reconstruct the extracted speech signal back in the time domain (the so-called permutation problem), how to deal with the case of more than two speakers in the room and when the speakers are moving. Our approach in this work is to use additional visual information to overcome these problems. We therefore wish to equip an intelligent office within the Centre of Digital Signal Processing at the Cardiff School of Engineering with microphones and cameras together with the necessary computing facilities to record examples of audio and visual signals for testing, initially with two well positioned speakers uttering distinct sounds, such as vowels and consenants, and then moving onto more speakers and movement; ultimately, recording natural continuous speech. The overall goal is to be able to demonstrate the ability to separate any one of the speakers utterances within the intelligent office which would then facilitate interaction, for example, with a voice recogniser or third party, at a remote location, as in teleconferencing.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1049/iet-ipr.2009.0042
发表时间: 2010-12-01
期刊: IET IMAGE PROCESSING
影响因子: 2.3
作者: [Aubrey, A. J., Hicks, Y. A., Chambers, J. A.]
通讯作者: Chambers, J. A.
DOI: --
发表时间: 2008-08
期刊: 2008 16th European Signal Processing Conference
影响因子: --
作者: [S. M. Naqvi;Y. Zhang;T. Tsalaile;S. Sanei;J. Chambers]
通讯作者: S. M. Naqvi;Y. Zhang;T. Tsalaile;S. Sanei;J. Chambers
DOI: 10.1109/jstsp.2010.2057198
发表时间: 2010
期刊: IEEE Journal of Selected Topics in Signal Processing
影响因子: 7.5
作者: [Naqvi S]
通讯作者: Naqvi S
Communications Signal Processing Based Solutions for Massive Machine-to-Machine Networks (M3NETs)
  • 批准号:
    EP/R006377/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $35.11万
  • 财政年份:
    2018
  • 负责人:
    Jonathon Chambers
  • 依托单位:
Signal Processing Solutions for the Networked Battlespace
  • 批准号:
    EP/K014307/2
  • 项目类别:
    Research Grant
  • 资助金额:
    $274.04万
  • 财政年份:
    2015
  • 负责人:
    Jonathon Chambers
  • 依托单位:
Signal Processing Solutions for the Networked Battlespace
  • 批准号:
    EP/K014307/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $464.65万
  • 财政年份:
    2013
  • 负责人:
    Jonathon Chambers
  • 依托单位:
Audio and Video Based Speech Separation for Multiple Moving Sources Within a Room Environment
  • 批准号:
    EP/H049665/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $38.32万
  • 财政年份:
    2010
  • 负责人:
    Jonathon Chambers
  • 依托单位:
海外基金