Explore new approaches to distant microphone speech recognition that combine information across multiple microphone array devices
Explore new approaches to distant microphone speech recognition that combine information across multiple microphone array devices
批准号:
2112956
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
使用语音与数字设备进行交流正变得越来越普遍。在过去的几年里,谷歌Home和亚马逊Alexa等设备已经进入了数百万家庭。让语音识别在家庭环境中正常工作是非常具有挑战性的。家里通常是一个非常嘈杂的地方,例如,如果设备放在厨房里,洗衣机可能正在运行,人们可能会在背景中说话。此外,说话的人通常离设备有几米远(“远距离麦克风”场景)。这是一个问题,因为语音信号很容易被其他可能更靠近麦克风的声源所控制。本项目将为远距离麦克风语音识别问题开发新颖的解决方案。该研究将由语言和听力研究小组在Jon Barker教授的指导下进行。它将利用Barker教授的研究小组在谷歌(http://spandh.dcs.shef.ac.uk/chime_challenge/)的支持下获得的新数据集('CHiME-5')。CHiME-5是一组真实家庭聚会的录音。数据由多个记录设备捕获,每个记录设备捕获视频和四个同步麦克风通道。这种独特的数据提供了一个机会,以解决新的研究问题,谎言在当前的语音技术范围之外。两个重点研究方向将被优先考虑,视觉驱动的波束形成算法:远距离麦克风语音识别最成功的方法是使用多个麦克风,并应用技术,增强来自某些方向的信号,同时抑制来自其他方向的信号。这需要检测和跟踪哪些方向是重要的。该项目将研究如何从视频信号中提取这些信息(例如,使用人员跟踪技术)。使用多个麦克风阵列进行语音识别:上述“波束成形”需要同步麦克风,它们彼此之间的位置已知。因此,它可以很容易地应用于相对位置不确定的多个设备(例如,在同一房间中组合两个谷歌home的输出)。CHiME-5数据在同一声学区域内有多达6个设备,因此为寻找解决该问题的新方案提供了独特的机会。我们可以从探索对独立识别系统的输出进行加权和融合的技术开始。语音识别系统已经发展成为非常复杂的软件。幸运的是,语音研究已经有效地开源了,社区现在集中在Kaldi语音识别工具包上。CHiME-5数据集将与开源Kaldi“基线”一起发布,该基线将代表单设备音频系统的最先进系统。它还将为培训系统提供一套“规则”,允许研究小组之间进行公平比较。这将为比较视听和多设备扩展的性能提供一个可靠的参考。这项研究将需要采用多种方法:视频人脸和人跟踪和波束形成算法;语音识别融合策略和信号质量评估技术。此外,有必要对基线识别器中采用的最先进技术有更全面的了解,包括卷积神经网络、i向量分析、说话人自适应训练、神经网络语言建模等。幸运的是,有许多优秀的教科书、指导论文和评论论文涵盖了这些领域。CHiME-5是一项复杂的“对话式”语音识别任务。训练和测试识别系统将需要大量的计算。现代语音识别器使用“深度学习”,这需要专业的GPU硬件。
英文摘要
It is becoming common for speech to be used to communicate with digital devices. In the last few years, devices such as Google Home and Amazon Alexa have arrived in millions of homes. Getting speech recognition to work well in home environments is very challenging. The home is often a very noisy place, for example, if the device is placed in the kitchen, the washing machine may be running and people could be talking in the background. Also, the person speaking is often several metres away from the device (the 'distant microphone' scenario). This is a problem because the speech signal may easily be dominated by other sound sources which may be closer to the microphones.This project will develop novel solutions to the distant microphone speech recognition problem. It will be conducted within the Speech and Hearing Research Group under the supervision of Prof. Jon Barker. It will take advantage of a new data set ('CHiME-5') that has been acquired by Prof. Barker's research team with support from Google (http://spandh.dcs.shef.ac.uk/chime_challenge/). CHiME-5 is a set of recordings of parties taking place in real homes. The data is captured with multiple recording devices, each of which captures video and four synchronised microphone channels. This unique data provides an opportunity to address new research questions lying outside the scope of current speech technology.Research questionsTwo key research directions will be prioritised,Visually-driven beamforming algorithms: The most successful approach to distant microphone speech recognition is to use multiple microphones and apply techniques that enhance the signals coming from some directions while suppressing the signals coming from others. This requires detecting and tracking which directions are important. The project will look at how this information might be extracted from the video signal (e.g.,using person tracking techniques.)Speech recognition with multiple microphone arrays: The 'beamforming' described above requires synchronised microphones with known positions with respect to each other. It can therefore be easily applied across multiple devices whose relative location is uncertain (e.g., combining outputs of two Google Homes in the same room). The CHiME-5 data has up to six devices within the same acoustic area and therefore provides a unique opportunity to find new solutions to this problem. A starting place would be to explore techniques for weighting and fusing the outputs of independent recognition systems.MethodologySpeech recognition systems have evolved into hugely complex pieces of software. Fortunately, speech research has been effectively open-sourced with the community now focused around the Kaldi speech recognition toolkit. The CHiME-5 data set will be published with an open-source Kaldi 'baseline' that will represent a state-of-the-art system for single device audio-only system. It will also provide a set of 'rules' for training systems that allowsfair comparison between research groups. This will provide a robust reference against which to compare the performance of audio-visual and multi-device extensions.The research will require a mixture of methods to be employed: video face and person tracking and beamforming algorithms; speech recognition fusion strategies, and signal quality assessment techniques. In addition, it will be necessary to have a fuller understanding of state-of-the-art techniques employed in the baseline recogniser, including convolutional neural networks, i-vector analysis, speaker-adaptive training, neural network language modelling, etc. Fortunately there are many excellent textbooks, tutorial papers and review papers that coverthese areas.CHiME-5 is a complex 'conversational' speech recognition task. Training and testing the recognition systems will be computationally demanding. Modern speech recognisers use 'deep learning' which requires specialist GPU hardware.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
脊髓新鉴定SNAPR神经元相关环路介导SCS电刺激抑制恶性瘙痒
-
批准号:82371478
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:焦英甫
-
依托单位:
tau轻子衰变与新物理模型唯象研究
-
批准号:11005033
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2010
-
负责人:李文君
-
依托单位:
HIV gp41的NHR区新靶点的确证及高效干预
-
批准号:81072676
-
项目类别:面上项目
-
资助金额:33.0万元
-
批准年份:2010
-
负责人:戴秋云
-
依托单位:
强子对撞机上新物理信号的多轻子末态研究
-
批准号:10675110
-
项目类别:面上项目
-
资助金额:36.0万元
-
批准年份:2006
-
负责人:蒋一
-
依托单位: