Nonparametric speaker recognition method using Earth Mover's Distance

Nonparametric speaker recognition method using Earth Mover's Distance
复制标题

DOI:
10.1093/ietisy/e89-d.3.1074
复制
发表时间:
2006-03-01
影响因子:
0.7
通讯作者:
Ren, F
Ren, F
中科院分区:
计算机科学4区
文献类型:
--
作者:
Kuroiwa, S;Umeda, Y;Ren, F

文献摘要

被引文献

相似文献

本文提出了一种基于非参数说话人模型和地球移动距离(EMD)的分布式说话人识别方法。在分布式说话人识别中,量化的特征向量被发送到服务器。说话人识别的传统方法--高斯混合模型(GMM)采用最大似然方法进行训练。然而,对量化后的数据进行连续密度函数的拟合比较困难。为了克服这一问题,该方法将每个说话人模型表示为由注册特征向量设计的依赖于说话人的VQ码直方图,并直接计算说话人模型的直方图与测试量化特征向量之间的距离。为了测量每个说话人模型与测试数据之间的距离,我们使用EMD来计算不同区段的直方图之间的距离。利用该方法进行了与文本无关的说话人辨认实验。与传统的GMM方法相比,该方法对量化数据的相对误差降低了32%。
In this paper, we propose a distributed speaker recognition method using a nonparametric speaker model and Earth Mover's Distance (EMD). In distributed speaker recognition, the quantized feature vectors are sent to a server. The Gaussian mixture model (GMM), the traditional method used for speaker recognition, is trained using the maximum likelihood approach. However, it is difficult to fit continuous density functions to quantized data. To overcome this problem, the proposed method represents each speaker model with a speaker-dependent VQ code histogram designed by registered feature vectors and directly calculates the distance between the histograms of speaker models and testing quantized feature vectors. To measure the distance between each speaker model and testing data, we use EMD which can calculate the distance between histograms with different bins. We conducted text-independent speaker identification experiments using the proposed method. Compared to results using the traditional GMM, the proposed method yielded relative error reductions of 32% for quantized data.