A STICKY HDP-HMM WITH APPLICATION TO SPEAKER DIARIZATION

A STICKY HDP-HMM WITH APPLICATION TO SPEAKER DIARIZATION
复制标题

DOI:
10.1214/10-aoas395
复制
发表时间:
2011-06-01
影响因子:
1.8
通讯作者:
Willsky, Alan S.
Willsky, Alan S.
中科院分区:
数学4区
文献类型:
--
作者:
Fox, Emily B.;Sudderth, Erik B.;Willsky, Alan S.

文献摘要

被引文献

相似文献

我们考虑的问题,扬声器日记,分割成对应于个别扬声器的时间段的会议的音频记录的问题。由于我们不被允许假定知道参加会议的人数,这个问题变得特别困难。为了解决这个问题,我们采取贝叶斯非参数方法来说话人日记化,该方法建立在Teh等人的分层狄利克雷过程隐马尔可夫模型(HDP-HMM)上。101(2006)1566-1581]。虽然基本的HDP-HMM往往过度分割的音频数据,创建冗余的状态,并在它们之间快速切换,我们描述了一个增强的HDP-HMM,提供有效的控制切换速率。我们还表明,这种增强使得有可能处理排放分布nonparametrically。为了将所得到的架构扩展到现实的日记化问题,我们开发了一种采样算法,该算法采用截断近似的Dirichlet过程来联合重新采样完整的状态序列,大大提高了混合率。与基准NIST数据集的工作,我们表明,我们的贝叶斯非参数架构产生国家的最先进的扬声器日记化的结果。
We consider the problem of speaker diarization, the problem of segmenting an audio recording of a meeting into temporal segments corresponding to individual speakers. The problem is rendered particularly difficult by the fact that we are not allowed to assume knowledge of the number of people participating in the meeting. To address this problem, we take a Bayesian nonparametric approach to speaker diarization that builds on the hierarchical Dirichlet process hidden Markov model (HDP-HMM) of Teh et al. [J. Amer. Statist. Assoc. 101 (2006) 1566-1581]. Although the basic HDP-HMM tends to over-segment the audio data-creating redundant states and rapidly switching among them-we describe an augmented HDP-HMM that provides effective control over the switching rate. We also show that this augmentation makes it possible to treat emission distributions nonparametrically. To scale the resulting architecture to realistic diarization problems, we develop a sampling algorithm that employs a truncated approximation of the Dirichlet process to jointly resample the full state sequence, greatly improving mixing rates. Working with a benchmark NIST data set, we show that our Bayesian nonparametric architecture yields state-of-the-art speaker diarization results.