Improved i-Vector Representation for Speaker Diarization

Improved i-Vector Representation for Speaker Diarization
复制标题

DOI:
10.1007/s00034-015-0206-2
复制
发表时间:
2015-12
期刊:
Circuits, Systems, and Signal Processing
影响因子:
--
通讯作者:
Yan Xu;I. Mcloughlin;Yan Song;Kui Wu
Yan Xu;I. Mcloughlin;Yan Song;Kui Wu
中科院分区:
其他
文献类型:
--
作者:
Yan Xu;I. Mcloughlin;Yan Song;Kui Wu

文献摘要

被引文献

相似文献

本文提出使用先前经过良好训练的深度神经网络(DNN)来增强用于说话者日记化的i向量表示。实际上,我们将通常用于训练通用背景模型(UBM)的高斯混合模型替换为使用不同大规模数据集训练的DNN。为了训练T矩阵,我们使用从DNN获得的监督UBM,使用滤波器组输入特征来计算后验信息,然后使用MFCC特征来训练UBM,而不是传统的从单个特征导出的无监督UBM。接下来,我们联合使用DNN和MFCC特征来计算零阶和一阶Baum-Welch统计量,以训练提取器,从中获得i向量。该系统将被证明是实现显着改善NIST 2008年说话人识别评估电话数据任务相比,国家的最先进的方法。
This paper proposes using a previously well-trained deep neural network (DNN) to enhance the i-vector representation used for speaker diarization. In effect, we replace the Gaussian mixture model typically used to train a universal background model (UBM), with a DNN that has been trained using a different large-scale dataset. To train theT-matrix, we use a supervised UBM obtained from the DNN using filterbank input features to calculate the posterior information and then MFCC features to train the UBM instead of a traditional unsupervised UBM derived from single features. Next we jointly use DNN and MFCC features to calculate the zeroth- and first-order Baum–Welch statistics for training an extractor from which we obtain the i-vector. The system will be shown to achieve a significant improvement on the NIST 2008 speaker recognition evaluation telephone data task compared to state-of-the-art approaches.