课题基金 / 基金详情

Representations of Speech Dynamics as Features for Speaker Recognition

Representations of Speech Dynamics as Features for Speaker Recognition
语音动力学的表示作为说话人识别的特征
批准号:
105523-2012
负责人:
Kenny, Patrick
金额:
$1.53万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2015
资助国家:
加拿大
项目状态:
已结题
起止时间:
2015-01-01 至 2016-12-31

项目摘要

项目成果

Kenny, Patrick的其他基金

相似基金

相关文献

中文摘要
翻译
近年来,说话人识别领域取得了非常迅速的进展。 自2005年以来,NIST说话人识别评估中的错误率下降了2- 3倍,2010年对“人类辅助说话人识别”的试点评估表明,NIST定义的核心说话人识别任务的机器性能与人类专家相当。 这一进展在很大程度上是由于成功的方法,如本征信道建模和联合因子分析(JFA)在衰减会话间变异性的问题。 这些方法使人们有可能处理两个话语之间的差异是否可以通过假设不同的扬声器或其他因素,如通道效应更好地解释的基本问题。 通过使用i向量(本质上是JFA模型中的隐藏变量)作为特征,该技术已经大大简化,这使得设计非常轻的重量,高性能分类器成为可能。 值得注意的是,这一进展是使用高斯混合模型(GARCH)作为语音的基本生成模型来实现的。 Gynecology可以模拟说话者声音的音色,但他们无法提取任何有关语音动态的信息,这些信息可能有助于区分说话者。 早期尝试使用各种类型的细分模型来解决这个问题,但这些方法都无法与甘精胰岛素相媲美,近年来这些方法只是偶尔被重新审视。 我建议解决这个问题,使用新的生成模型的语音动力学设计这样一种方式,他们可以用来提取鲍姆-韦尔奇样的统计数据从语音数据。 这将允许与我已经开发的i-vector/PLDA设备轻松集成。
英文摘要
There has been very rapid progress in the field of speaker recognition in recent years. Error rates in the NIST speaker recognition evaluations have decreased by a factor of 2--3 since 2005 and a pilot evaluation of "human assisted speaker recognition" in 2010 indicated that machine performance on the core speaker recognition task as defined by NIST is comparable to that of human experts. This progress is largely due to the success of methods such as eigenchannel modeling and Joint Factor Analysis (JFA) in attenuating the problem of intersession variability. These methods made it possible to get a handle on the fundamental question of whether the difference between two utterances could be better accounted for by postulating different speakers or other factors such as channel effects. The technology has been greatly simplified by using i-vectors (essentially the hidden variables in the JFA model) as features, a development which made it possible to design very light weight, high performance classifiers. It is remarkable that this progress has been achieved using Gaussian mixture models (GMMs) as the underlying generative model of speech. GMMs can model the timbre of a speaker's voice but they are incapable of extracting any information about speech dynamics which might be useful for distinguishing between speakers. Early attempts to tackle this problem using various types of segment models were no match for GMMs and these approaches have only been sporadically revisited in recent years. I propose to address this problem using new generative models of speech dynamics designed in such a way that they can be used to extract Baum-Welch-like statistics from speech data. This will allow for easy integration with the i-vector/PLDA apparatus that I have already developed.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
JFA for text-dependent speaker verification
Representations of Speech Dynamics as Features for Speaker Recognition
JFA for text-dependent speaker verification
Representations of Speech Dynamics as Features for Speaker Recognition
海外基金