Strategies to Improve the Robustness of Agglomerative Hierarchical Clustering Under Data Source Variation for Speaker Diarization

Strategies to Improve the Robustness of Agglomerative Hierarchical Clustering Under Data Source Variation for Speaker Diarization
复制标题

DOI:
10.1109/tasl.2008.2002085
复制
发表时间:
2008-11-01
影响因子:
--
通讯作者:
Narayanan, Shrikanth S.
Narayanan, Shrikanth S.
中科院分区:
其他
文献类型:
--
作者:
Han, Kyu J.;Kim, Samuel;Narayanan, Shrikanth S.

文献摘要

被引文献

相似文献

由于聚类分层聚类的处理结构简单、性能可接受,目前许多先进的说话人聚类系统都采用聚类分层聚类作为说话人聚类策略。然而,AHC在数据源变化情况下存在性能鲁棒性问题。在本文中,我们解决了这个问题。重点研究了基于贝叶斯信息准则(BIC)的聚类停止方法和基于广义似然比(GLR)的合并聚类选择方案。首先,提出了一种基于信息变化率(ICR)的AHC替代停止方法。通过在多个会议语料库上的实验,证明了该方法对数据源变化的鲁棒性优于基于bic的方法。采用该方法对数字错误率(DER)的平均改善为8.76%(绝对)或35.77%(相对)。本文还引入了一种选择性AHC (SAHC),它首先使用基于icr的停止方法对长度大于3秒的语音片段运行AHC,然后将较短的语音片段分类到初始AHC给出的一个聚类中。AHC的修改版本源于我们之前的分析,即数据源中短语音回合(或片段)的比例是导致基于glr的合并聚类选择方案中出现鲁棒性问题的重要因素。SAHC获得的额外性能改进在平均DER方面为3.45%(绝对)或14.08%(相对)。
Many current state-of-the-art speaker diarization systems exploit agglomerative hierarchical clustering (AHC) as their speaker clustering strategy, due to its simple processing structure and acceptable level of performance. However, AHC is known to suffer from performance robustness under data source variation. In this paper, we address this problem. We specifically focus on the issues associated with the widely used clustering stopping method based on Bayesian information criterion (BIC) and the merging-cluster selection scheme based on generalized likelihood ratio (GLR). First, we propose a novel alternative stopping method for AHC based on information change rate (ICR). Through experiments on several meeting corpora, the proposed method is demonstrated to be more robust to data source variation than the BIC-based one. The average improvement obtained in diarization error rate (DER) by this method is 8.76% (absolute) or 35.77% (relative). We also introduce a selective AHC (SAHC) in the paper, which first runs AHC with the ICR-based stopping method only on speech segments longer than 3 s and then classifies shorter speech segments into one of the clusters given by the initial AHC. This modified version of AHC is motivated by our previous analysis that the proportion of short speech turns (or segments) in a data source is a significant factor contributing to the robustness problem arising in the GLR-based merging-cluster selection scheme. The additional performance improvement obtained by SAHC is 3.45% (absolute) or 14.08% (relative) in terms of averaged DER.