Pairwise Discriminative Speaker Verification in the ${\rm I}$-Vector Space

Pairwise Discriminative Speaker Verification in the ${\rm I}$-Vector Space
复制标题

DOI:
10.1109/tasl.2013.2245655
复制
发表时间:
2011-05
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Sandro Cumani;N. Brümmer;L. Burget;P. Laface;Oldrich Plchot;Vasileios Vasilakakis
Sandro Cumani;N. Brümmer;L. Burget;P. Laface;Oldrich Plchot;Vasileios Vasilakakis
中科院分区:
其他
文献类型:
--
作者:
Sandro Cumani;N. Brümmer;L. Burget;P. Laface;Oldrich Plchot;Vasileios Vasilakakis

文献摘要

被引文献

相似文献

这项工作提出了一种新的和有效的方法来区分说话人确认的i-向量空间。我们说明了一个线性判别分类器的发展,该分类器被训练来区分假设,即在一次试验中的一对特征向量属于同一发言人或不同的发言人。这种方法替代了通常的区分说话者和所有其他说话者的区分设置。我们使用一个判别分类器的基础上的支持向量机(SVM)的训练,估计一个对称的二次函数的参数近似的对数似然比得分没有明确的建模的i-向量分布的生成概率线性判别分析(PLDA)模型。训练这些模型是可行的,因为不需要扩展i-向量对,即使对于中等大小的训练集,扩展i-向量对也是昂贵的或者甚至是不可能的。在NIST 2010说话人识别评估的tel-tel扩展核心条件上进行的实验结果在归一化检测成本函数和等错误率方面与生成模型所获得的结果具有竞争力。此外,我们证明了可以训练一个与性别无关的判别模型,该模型可以达到最先进的准确性,与性别相关系统的准确性相当,从而在训练和测试中节省内存和执行时间。
This work presents a new and efficient approach to discriminative speaker verification in the i-vector space. We illustrate the development of a linear discriminative classifier that is trained to discriminate between the hypothesis that a pair of feature vectors in a trial belong to the same speaker or to different speakers. This approach is alternative to the usual discriminative setup that discriminates between a speaker and all the other speakers. We use a discriminative classifier based on a Support Vector Machine (SVM) that is trained to estimate the parameters of a symmetric quadratic function approximating a log-likelihood ratio score without explicit modeling of the i-vector distributions as in the generative Probabilistic Linear Discriminant Analysis (PLDA) models. Training these models is feasible because it is not necessary to expand the i -vector pairs, which would be expensive or even impossible even for medium sized training sets. The results of experiments performed on the tel-tel extended core condition of the NIST 2010 Speaker Recognition Evaluation are competitive with the ones obtained by generative models, in terms of normalized Detection Cost Function and Equal Error Rate. Moreover, we show that it is possible to train a gender-independent discriminative model that achieves state-of-the-art accuracy, comparable to the one of a gender-dependent system, saving memory and execution time both in training and in testing.