Speaker Verification Using Sparse Representations on Total Variability i-vectors

Speaker Verification Using Sparse Representations on Total Variability i-vectors
复制标题

DOI:
10.21437/interspeech.2011-149
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
Ming Li;Xiang Zhang;Yonghong Yan;Shrikanth S. Narayanan
Ming Li;Xiang Zhang;Yonghong Yan;Shrikanth S. Narayanan
中科院分区:
其他
文献类型:
--
作者:
Ming Li;Xiang Zhang;Yonghong Yan;Shrikanth S. Narayanan

文献摘要

被引文献

相似文献

本文在进行类内协方差归一化和线性判别分析通道补偿后,利用二次约束最小化计算得到的稀疏表示对低维总变异性空间中的i向量进行建模。首先,我们提出了背景归一化残差作为评分准则。其次,我们证明了利用Tnorm数据作为过完备字典中的非目标样本可以有效地实现Tnorm。最后,通过融合传统的基于i向量的支持向量机(SVM)和余弦距离评分系统,我们证明了系统整体性能的提高。实验结果表明,该融合系统在NIST SRE 2008的单-单多语言手持电话任务上,经Tnorm处理后的平均错误率(EER)分别为4.05%(男性)和5.25%(女性),在男性和女性任务上的相对错误率分别降低7.1%和4.9%,优于支持向量机基线。
In this paper, the sparse representation computed by lminimization with quadratic constraints is employed to model the i-vectors in the low dimensional total variability space after performing the Within-Class Covariance Normalization and Linear Discriminate Analysis channel compensation. First, we propose the background normalized l residual as a scoring criterion. Second, we demonstrate that the Tnorm can be efficiently achieved by using the Tnorm data as the non-target samples in the over-complete dictionary. Finally, by fusing with the conventional i-vector based support vector machine (SVM) and cosine distance scoring system, we demonstrate overall system performance improvement. Experimental results show that the proposed fusion system achieved 4.05% (male) and 5.25% (female) equal error rate (EER) after Tnorm on the single-single multi-language handheld telephone task of NIST SRE 2008 and outperformed the SVM baseline by yielding 7.1% and 4.9% relative EER reduction for the male and female tasks, respectively.