Audio and Video Based Speech Separation for Multiple Moving Sources Within a Room Environment
Audio and Video Based Speech Separation for Multiple Moving Sources Within a Room Environment
批准号:
EP/H049665/1
负责人:
Jonathon Chambers
金额:
$38.32万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Human beings have developed a unique ability to communicate within a noisy environment, such as at a cocktail party. This skill is dependent upon the use of both the aural and visual senses together with sophisticated processing within the brain. To mimic this ability within a machine is very challenging, particularly if the humans are moving, such as in a teleconferencing context, when human speakers are walking around a room. In the field of signal processing researchers have developed techniques to separate one speech signal from a mixture of such signals, as would be measured by a number of microphones, on the basis of only audio information with the assumption that the humans are static and typically no more than two humans are within the room. Such approaches have generally been found to fail, however, when the human speakers are moving and when there are more than two in number. Fundamentally new approaches are therefore necessary to advance the state-of-the-art in the field. Professor Chambers and his team at Loughborough University were the first in the UK to propose a new approach on the basis of combined audio and video processing to solve the source separation problem, but their preliminary approach identified major challenges in audio-visual speaker localization, tracking and separation which must be solved to provide a practical solution for speech separation for multiple moving sources within a room environment. These findings motivate this new project in which world-leading teams at the University of Surrey, led by Professor Kittler, and at the GIPSA Lab, Grenoble, France, headed by Professor Jutten, are ready to work with Professor Chambers and his team at Loughborough University to advance the state-of-the-art in the field.In this new project, two postdoctoral researchers will be employed, one at Loughborough and another at Surrey. The first will focus on the development of fundamentally new speech source separation algorithms for moving speakers by using geometrical room acoustic (for example location and number of sources, descriptions of their movement) information provided by the second researcher. The research team at Grenoble will provide technical guidance on the basis of their considerable experience in source separation throughout the project and will work on providing an acoustic noise model for the room environment which will also aid the speech separation process. To achieve these tasks, frequency domain based beamforming algorithms will be developed which exploit microphone arrays having more microphones than speakers so that new data independent superdirective robust beamformer design methods can be exploited using mathematical convex optimization. Additionally, further geometic information will be exploited to introduce robustness to errors in the localization information describing the desired source and the interference. To improve the localization information an array of collaborative cameras will be used and both audio and visual information will be used. Advanced methods from particle filtering and probabilistic data association will be exploited for improving the tracking performance. Finally, visual voice activity detection will be used to determine the active sources within the beamforming operations. We emphasize that this work is not implementation-driven, so computational complexity for real-time realization will not be a focus; this would be the subject of a future project.All of the new algorithms will be evaluated both in terms of objective and subjective performance measures on labelled audio and visual datasets acquired at Loughbourgh and Surrey, and from the CHIL seminar room at the Karlsruhe University (UKA), Germany. To ensure this pioneering work has maximum impact on the UK and international academic and research communities all the algorithms and datasets will be made available through the project website.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Fast pose invariant face recognition using super coupled multiresolution Markov Random Fields on a GPU
在 GPU 上使用超耦合多分辨率马尔可夫随机场进行快速姿势不变人脸识别
DOI:
10.1016/j.patrec.2014.05.017
发表时间:
2014
期刊:
Pattern Recognition Letters
影响因子:
5.1
作者:
[Rahimzadeh Arashloo S]
通讯作者:
Rahimzadeh Arashloo S
DOI:
--
发表时间:
2011-08
期刊:
2011 19th European Signal Processing Conference
影响因子:
--
作者:
[S. M. Naqvi;Muhammad Salman Khan;Qingju Liu;Wenwu Wang;J. Chambers]
通讯作者:
S. M. Naqvi;Muhammad Salman Khan;Qingju Liu;Wenwu Wang;J. Chambers
Robust Feature Selection for Scaling Ambiguity Reduction in Audio-Visual Convolutive BSS
用于音视频卷积 BSS 中缩放模糊度减少的鲁棒特征选择
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[Syed Mohsen Naqvi (Author)]
通讯作者:
Syed Mohsen Naqvi (Author)
DOI:
10.1109/jstsp.2010.2057198
发表时间:
2010
期刊:
IEEE Journal of Selected Topics in Signal Processing
影响因子:
7.5
作者:
[Naqvi S]
通讯作者:
Naqvi S
DOI:
10.1186/1687-6180-2012-183
发表时间:
2012-08
期刊:
EURASIP Journal on Advances in Signal Processing
影响因子:
1.9
作者:
[Yanfeng Liang;S. M. Naqvi;J. Chambers]
通讯作者:
Yanfeng Liang;S. M. Naqvi;J. Chambers
共 6 条
Communications Signal Processing Based Solutions for Massive Machine-to-Machine Networks (M3NETs)
-
批准号:EP/R006377/1
-
项目类别:Research Grant
-
资助金额:$35.11万
-
财政年份:2018
-
负责人:Jonathon Chambers
-
依托单位:
Signal Processing Solutions for the Networked Battlespace
-
批准号:EP/K014307/2
-
项目类别:Research Grant
-
资助金额:$274.04万
-
财政年份:2015
-
负责人:Jonathon Chambers
-
依托单位:
Signal Processing Solutions for the Networked Battlespace
-
批准号:EP/K014307/1
-
项目类别:Research Grant
-
资助金额:$464.65万
-
财政年份:2013
-
负责人:Jonathon Chambers
-
依托单位:
Novel Communications Signal Processing Techs. for Transmission Over MIMO Frequency Selective Wireless Channels Using Polynomial Matrix Decompositions
-
批准号:EP/F065477/1
-
项目类别:Research Grant
-
资助金额:$48.03万
-
财政年份:2008
-
负责人:Jonathon Chambers
-
依托单位:
Multi-Modal Blind Source Separation Algorithms
-
批准号:EP/C535308/2
-
项目类别:Research Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Jonathon Chambers
-
依托单位:
海外基金