Analysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priors

Analysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priors
复制标题

DOI:
10.1109/taslp.2019.2955293
复制
发表时间:
2020-01-01
影响因子:
5.4
通讯作者:
Cernocky, Jan
Cernocky, Jan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Diez, Mireia;Burget, Lukas;Cernocky, Jan

文献摘要

被引文献

相似文献

在我们以前的工作中,我们介绍了我们的贝叶斯隐马尔可夫模型与本征语音先验,这已被认为是最先进的模型发言人日记。在这篇文章中,我们提出了一个更完整的分析日记系统。文中详细描述了模型的推导过程,并给出了所有更新公式的推导过程,使读者对算法有一个完整的理解。对模型参数的影响、敏感性和交互作用进行了分析,为模型参数的优化设置提供了依据。新引入的说话人正则化系数允许我们控制话语中推断的说话人的数量。一个朴素说话人模型合并策略,它允许驱动变分推理出局部最优。不同的日志化方案的实验CALLHOME和DIHARD数据集。
In our previous work, we introduced our Bayesian Hidden Markov Model with eigenvoice priors, which has been recently recognized as the state-of-the-art model for Speaker Diarization. In this article we present a more complete analysis of the Diarization system. The inference of the model is fully described and derivations of all update formulas are provided for a complete understanding of the algorithm. An extensive analysis on the effect, sensitivity and interactions of all model parameters is provided, which might be used as a guide for their optimal setting. The newly introduced speaker regularization coefficient allows us to control the number of speakers inferred in an utterance. A naive speaker model merging strategy is also presented, which allows to drive the variational inference out of local optima. Experiments for the different diarization scenarios are presented on CALLHOME and DIHARD datasets.