课题基金 / 基金详情

Representations of Speech Dynamics as Features for Speaker Recognition

Representations of Speech Dynamics as Features for Speaker Recognition
语音动力学的表示作为说话人识别的特征
批准号:
105523-2012
负责人:
Kenny, Patrick
金额:
$1.53万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2015
资助国家:
加拿大
项目状态:
已结题
起止时间:
2015-01-01 至 2016-12-31

项目摘要

项目成果

Kenny, Patrick的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
There has been very rapid progress in the field of speaker recognition in recent years. Error rates in the NIST speaker recognition evaluations have decreased by a factor of 2--3 since 2005 and a pilot evaluation of "human assisted speaker recognition" in 2010 indicated that machine performance on the core speaker recognition task as defined by NIST is comparable to that of human experts. This progress is largely due to the success of methods such as eigenchannel modeling and Joint Factor Analysis (JFA) in attenuating the problem of intersession variability. These methods made it possible to get a handle on the fundamental question of whether the difference between two utterances could be better accounted for by postulating different speakers or other factors such as channel effects. The technology has been greatly simplified by using i-vectors (essentially the hidden variables in the JFA model) as features, a development which made it possible to design very light weight, high performance classifiers. It is remarkable that this progress has been achieved using Gaussian mixture models (GMMs) as the underlying generative model of speech. GMMs can model the timbre of a speaker's voice but they are incapable of extracting any information about speech dynamics which might be useful for distinguishing between speakers. Early attempts to tackle this problem using various types of segment models were no match for GMMs and these approaches have only been sporadically revisited in recent years. I propose to address this problem using new generative models of speech dynamics designed in such a way that they can be used to extract Baum-Welch-like statistics from speech data. This will allow for easy integration with the i-vector/PLDA apparatus that I have already developed.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
JFA for text-dependent speaker verification
Representations of Speech Dynamics as Features for Speaker Recognition
JFA for text-dependent speaker verification
Representations of Speech Dynamics as Features for Speaker Recognition
海外基金