A new kernel for SVM MLLR based speaker recognition

A new kernel for SVM MLLR based speaker recognition
复制标题

DOI:
10.21437/interspeech.2007-130
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
Z. Karam;W. Campbell
Z. Karam;W. Campbell
中科院分区:
其他
文献类型:
--
作者:
Z. Karam;W. Campbell

文献摘要

被引文献

相似文献

说话人识别使用支持向量机(SVMs)的特征来自生成模型已被证明是表现良好的。通常,通用背景模型(UBM)适于每个话语,产生在SVM中使用的一组特征。我们考虑的情况下,UBM是高斯混合模型(GMM),最大似然线性回归(MLLR)适应是用来适应的UBM的手段。我们研究了两种可能的SVM特征扩展,在这种情况下出现:第一,GMM超向量是通过堆叠的适应GMM的手段,第二个由MLLR变换的元素。我们研究了与这些扩展相关的几个内核。我们表明,这两个扩展是等价的内核的适当选择。在NIST SRE 2006语料库上进行的实验清楚地表明,我们选择的内核,这是由GARCH之间的距离度量,优于ad-hoc的。我们还应用SVM滋扰属性投影(NAP)的内核作为一种形式的通道补偿,并表明,与一个适当的选择内核,我们实现的结果相比,现有的基于SVM的识别。索引术语:说话人识别,MLLR,SVM,超向量
Speaker recognition using support vector machines (SVMs) with features derived from generative models has been shown to perform well. Typically, a universal background model (UBM) is adapted to each utterance yielding a set of features that are used in an SVM. We consider the case where the UBM is a Gaussian mixture model (GMM), and maximum likelihood linear regression (MLLR) adaptation is used to adapt the means of the UBM. We examine two possible SVM feature expansions that arise in this context: the first, a GMM supervector is constructed by stacking the means of the adapted GMM, and the second consists of the elements of the MLLR transform. We examine several kernels associated with these expansions. We show that both expansions are equivalent given an appropriate choice of kernels. Experiments performed on the NIST SRE 2006 corpus clearly highlight that our choice of kernels, which are motivated by distance metrics between GMMs, outperform ad-hoc ones. We also apply SVM nuisance attribute projection (NAP) to the kernels as a form of channel compensation and show that, with a proper choice of kernel, we achieve results comparable to existing SVM based recognizers. Index Terms: speaker recognition, MLLR, SVM, supervector