A Study on Online Source Extraction in the Presence of Changing Speaker Positions

A Study on Online Source Extraction in the Presence of Changing Speaker Positions
复制标题

DOI:
10.1007/978-3-030-31372-2_17
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Jens Heitkaemper;T. Fehér;M. Freitag;Reinhold Häb-Umbach
Jens Heitkaemper;T. Fehér;M. Freitag;Reinhold Häb-Umbach
中科院分区:
其他
文献类型:
--
作者:
Jens Heitkaemper;T. Fehér;M. Freitag;Reinhold Häb-Umbach

文献摘要

被引文献

相似文献

多话语音和移动说话人仍然是自动语音识别系统面临的重大挑战。假设目标说话者的注册话语是可用的,最近提出了所谓的SpeakerBeam概念,从语音混合中提取目标说话者。如果有多通道输入,可以利用扬声器的空间特性来支持源提取。在这篇文章中,我们研究了利用这种空间信息的不同方法。特别是,我们感兴趣的问题是,如果目标说话人改变了他/她的位置,这些信息有多大用处?为此,我们提出了一种基于speakerbeam的源提取网络,该网络通过递归更新波束形成器系数来适应移动的扬声器。在两个数据集上给出了实验结果,一个是人工创建的房间脉冲响应,另一个是在会议室记录的真实房间脉冲响应和噪声。有趣的是,即使说话人的位置改变了,空间特征也被证明是有利的。
Multi-talker speech and moving speakers still pose a significant challenge to automatic speech recognition systems. Assuming an enrollment utterance of the target speakeris available, the so-called SpeakerBeam concept has been recently proposed to extract the target speaker from a speech mixture. If multi-channel input is available, spatial properties of the speaker can be exploited to support the source extraction. In this contribution we investigate different approaches to exploit such spatial information. In particular, we are interested in the question, how useful this information is if the target speaker changes his/her position. To this end, we present a SpeakerBeam-based source extraction network that is adapted to work on moving speakers by recursively updating the beamformer coefficients. Experimental results are presented on two data sets, one with artificially created room impulse responses, and one with real room impulse responses and noise recorded in a conference room. Interestingly, spatial features turn out to be advantageous even if the speaker position changes.