Speaker age estimation using i-vectors

Speaker age estimation using i-vectors
复制标题

DOI:
10.1016/j.engappai.2014.05.003
复制
发表时间:
2014-09-01
影响因子:
8
通讯作者:
van Leeuwen, David A.
van Leeuwen, David A.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bahari, Mohamad Hasan;McLaren, Mitchell;van Leeuwen, David A.

文献摘要

被引文献

相似文献

本文提出了一种基于i向量的语音信号年龄估计方法。在该方法中,每个话语由其相应的i向量建模。然后,采用类内协方差归一化技术对会话可变性进行补偿。最后,应用最小二乘支持向量回归(LSSVR)估计说话人的年龄。该方法在美国国家标准与技术研究院(NIST) 2010年和2008年说话人识别评估数据库的电话会话中进行了训练和测试。评估结果表明,与不同的传统方法相比,该方法的平均绝对误差显著降低,说话人实足年龄与估计说话人年龄之间的Pearson相关系数显著提高。得到的平均绝对误差和相关系数相对于我们最好的基线系统分别提高了5%和2%左右。最后,分析了影响年龄估计系统的主要因素,即话语长度和口语的影响。(C) 2014 Elsevier Ltd.版权所有。
In this paper, a new approach for age estimation from speech signals based on i-vectors is proposed. In this method, each utterance is modeled by its corresponding i-vector. Then, a Within-Class Covariance Normalization technique is used for session variability compensation. Finally, a least squares support vector regression (LSSVR) is applied to estimate the age of speakers. The proposed method is trained and tested on telephone conversations of the National Institute for Standard and Technology (NIST) 2010 and 2008 speaker recognition evaluation databases. Evaluation results show that the proposed method yields significantly lower mean absolute error and higher Pearson correlation coefficient between chronological speaker age and estimated speaker age compared to different conventional schemes. The obtained relative improvements of mean absolute error and correlation coefficient compared to our best baseline system are around 5% and 2% respectively. Finally, the effect of some major factors influencing the proposed age estimation system, namely utterance length and spoken language are analyzed. (C) 2014 Elsevier Ltd. All rights reserved.