An i-vector Extractor Suitable for Speaker Recognition with both Microphone and Telephone Speech

An i-vector Extractor Suitable for Speaker Recognition with both Microphone and Telephone Speech
复制标题

DOI:
--
复制
发表时间:
2010
影响因子:
--
通讯作者:
Mohammed Senoussaoui;P. Kenny;N. Dehak;P. Dumouchel
Mohammed Senoussaoui;P. Kenny;N. Dehak;P. Dumouchel
中科院分区:
--
文献类型:
--
作者:
Mohammed Senoussaoui;P. Kenny;N. Dehak;P. Dumouchel

文献摘要

被引文献

相似文献

人们普遍认为,当有足够的背景训练数据来处理传输信道的干扰效应时,说话人确认系统表现得更好。还已知的是,当训练数据的声音环境与使用环境(测试环境)的声音环境相似时,这些系统表现最佳。然而,对于某些应用,来自相同类型的声音环境的训练数据是稀缺的,而来自不同类型的环境的大量数据是可用的。在本文中,我们提出了一种新的架构,文本无关的说话人验证系统,令人满意的训练凭借有限数量的特定于应用程序的数据,补充了足够数量的训练数据,从一些其他方面。该架构基于Dehak [1]提出的从低维空间(全变率空间)提取参数(i向量)。我们的目标是扩展Dehak的工作稀疏数据,即麦克风语音的说话人识别。主要的挑战是克服这样一个事实,即没有足够的应用程序特定的数据来准确地估计总变异协方差矩阵。我们提出了一种基于联合因子分析(JFA)的方法来估计麦克风本征信道(稀疏数据)与电话本征信道(足够的数据)。对于分类,我们实验了以下两种方法:支持向量机(SVM)和余弦距离评分(CDS)分类器,基于余弦距离。我们目前的识别结果的一部分,女性的声音在采访数据的NIST 2008年SRE。当我们的系统与最先进的JFA融合时,可以获得最佳的性能。我们实现了13%的相对改善等错误率和检测成本函数的最小值从0.0219下降到0.0164。
It is widely believed that speaker verification systems perform better when there is sufficient background training data to deal with nuisance effects of transmission channels. It is also known that these systems perform at their best when the sound environment of the training data is similar to that of the context of use (test context). For some applications however, training data from the same type of sound environment is scarce, whereas a considerable amount of data from a different type of environment is available. In this paper, we propose a new architecture for text-independent speaker verification systems that are satisfactorily trained by virtue of a limited amount of application-specific data, supplemented with a sufficient amount of training data from some other context. This architecture is based on the extraction of parameters (i-vectors) from a low-dimensional space (total variability space) proposed by Dehak [1]. Our aim is to extend Dehak’s work to speaker recognition on sparse data, namely microphone speech. The main challenge is to overcome the fact that insufficient application-specific data is available to accurately estimate the total variability covariance matrix. We propose a method based on Joint Factor Analysis (JFA) to estimate microphone eigenchannels (sparse data) with telephone eigenchannels (sufficient data). For classification, we experimented with the following two approaches: Support Vector Machines (SVM) and Cosine Distance Scoring (CDS) classifier, based on cosine distances. We present recognition results for the part of female voices in the interview data of the NIST 2008 SRE. The best performance is obtained when our system is fused with the state-of-the-art JFA. We achieve 13% relative improvement on equal error rate and the minimum value of detection cost function decreases from 0.0219 to 0.0164.