Explore new approaches to distant microphone speech recognition that combine information across multiple microphone array devices
Explore new approaches to distant microphone speech recognition that combine information across multiple microphone array devices
批准号:
2112956
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
It is becoming common for speech to be used to communicate with digital devices. In the last few years, devices such as Google Home and Amazon Alexa have arrived in millions of homes. Getting speech recognition to work well in home environments is very challenging. The home is often a very noisy place, for example, if the device is placed in the kitchen, the washing machine may be running and people could be talking in the background. Also, the person speaking is often several metres away from the device (the 'distant microphone' scenario). This is a problem because the speech signal may easily be dominated by other sound sources which may be closer to the microphones.This project will develop novel solutions to the distant microphone speech recognition problem. It will be conducted within the Speech and Hearing Research Group under the supervision of Prof. Jon Barker. It will take advantage of a new data set ('CHiME-5') that has been acquired by Prof. Barker's research team with support from Google (http://spandh.dcs.shef.ac.uk/chime_challenge/). CHiME-5 is a set of recordings of parties taking place in real homes. The data is captured with multiple recording devices, each of which captures video and four synchronised microphone channels. This unique data provides an opportunity to address new research questions lying outside the scope of current speech technology.Research questionsTwo key research directions will be prioritised,Visually-driven beamforming algorithms: The most successful approach to distant microphone speech recognition is to use multiple microphones and apply techniques that enhance the signals coming from some directions while suppressing the signals coming from others. This requires detecting and tracking which directions are important. The project will look at how this information might be extracted from the video signal (e.g.,using person tracking techniques.)Speech recognition with multiple microphone arrays: The 'beamforming' described above requires synchronised microphones with known positions with respect to each other. It can therefore be easily applied across multiple devices whose relative location is uncertain (e.g., combining outputs of two Google Homes in the same room). The CHiME-5 data has up to six devices within the same acoustic area and therefore provides a unique opportunity to find new solutions to this problem. A starting place would be to explore techniques for weighting and fusing the outputs of independent recognition systems.MethodologySpeech recognition systems have evolved into hugely complex pieces of software. Fortunately, speech research has been effectively open-sourced with the community now focused around the Kaldi speech recognition toolkit. The CHiME-5 data set will be published with an open-source Kaldi 'baseline' that will represent a state-of-the-art system for single device audio-only system. It will also provide a set of 'rules' for training systems that allowsfair comparison between research groups. This will provide a robust reference against which to compare the performance of audio-visual and multi-device extensions.The research will require a mixture of methods to be employed: video face and person tracking and beamforming algorithms; speech recognition fusion strategies, and signal quality assessment techniques. In addition, it will be necessary to have a fuller understanding of state-of-the-art techniques employed in the baseline recogniser, including convolutional neural networks, i-vector analysis, speaker-adaptive training, neural network language modelling, etc. Fortunately there are many excellent textbooks, tutorial papers and review papers that coverthese areas.CHiME-5 is a complex 'conversational' speech recognition task. Training and testing the recognition systems will be computationally demanding. Modern speech recognisers use 'deep learning' which requires specialist GPU hardware.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
脊髓新鉴定SNAPR神经元相关环路介导SCS电刺激抑制恶性瘙痒
-
批准号:82371478
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:焦英甫
-
依托单位:
tau轻子衰变与新物理模型唯象研究
-
批准号:11005033
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2010
-
负责人:李文君
-
依托单位:
HIV gp41的NHR区新靶点的确证及高效干预
-
批准号:81072676
-
项目类别:面上项目
-
资助金额:33.0万元
-
批准年份:2010
-
负责人:戴秋云
-
依托单位:
强子对撞机上新物理信号的多轻子末态研究
-
批准号:10675110
-
项目类别:面上项目
-
资助金额:36.0万元
-
批准年份:2006
-
负责人:蒋一
-
依托单位: