An Initialization Scheme for Meeting Separation with Spatial Mixture Models

An Initialization Scheme for Meeting Separation with Spatial Mixture Models
复制标题

空间混合模型满足分离的初始化方案

DOI:
--
复制
发表时间:
2022
期刊:
Interspeech
影响因子:
--
通讯作者:
Reinhold Haeb
Reinhold Haeb
中科院分区:
--
文献类型:
--
作者:
Christoph Boeddeker;Tobias Cord;Thilo von Neumann;Reinhold Haeb

文献摘要

被引文献

相似文献

空间混合模型(SMM)支持的声学波束形成已被广泛用于同时活动说话人的分离。然而,几乎没有考虑将会议数据分开,因为会议数据的特点是录音很长,讲话只有部分重叠。在这一贡献中,我们证明了通常只有一个说话人是活跃的这一事实可以被用来巧妙地初始化采用时变类先验的SMM。在Libricss上的实验表明,与基于Dirichlet分布的随机初始化相比,本文提出的初始化方案在下行语音识别任务中获得了明显更低的误词率(WER)。只需知道说话人的数量,我们就获得了5.9%的WER,这与该数据集上报告的最好的WER相当。此外,从混合模型估计的说话人活动用作基于空间信息的二值化。
Spatial mixture model (SMM) supported acoustic beamforming has been extensively used for the separation of simultaneously active speakers. However, it has hardly been considered for the separation of meeting data, that are characterized by long recordings and only partially overlapping speech. In this contribution, we show that the fact that often only a single speaker is active can be utilized for a clever initialization of an SMM that employs time-varying class priors. In experiments on LibriCSS we show that the proposed initialization scheme achieves a significantly lower Word Error Rate (WER) on a downstream speech recognition task than a random initialization of the class probabilities by drawing from a Dirichlet distribution. With the only requirement that the number of speakers has to be known, we obtain a WER of 5.9 %, which is comparable to the best reported WER on this data set. Furthermore, the estimated speaker activity from the mixture model serves as a diarization based on spatial information.