ROBUST TEXT-INDEPENDENT SPEAKER IDENTIFICATION USING GAUSSIAN MIXTURE SPEAKER MODELS

ROBUST TEXT-INDEPENDENT SPEAKER IDENTIFICATION USING GAUSSIAN MIXTURE SPEAKER MODELS
复制标题

DOI:
10.1109/89.365379
复制
发表时间:
1995-01-01
期刊:
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
影响因子:
--
通讯作者:
ROSE, RC
ROSE, RC
中科院分区:
其他
文献类型:
--
作者:
REYNOLDS, DA;ROSE, RC

文献摘要

被引文献

相似文献

本文介绍并鼓励使用高斯混合模型(GMM)进行稳健的与文本无关的说话人识别,GMM的单个高斯分量表示一些一般的说话人相关的频谱形状,这些形状对于建模说话人身份是有效的。这项工作的重点是应用于需要使用来自不受限制的会话语音的短话语的高识别率和对电话信道上的传输产生的降级的鲁棒性的应用。在49个说话人对话电话语音库上对高斯混合说话人模型进行了完整的实验评估,实验检查了算法问题(初始化、方差限制、模型阶数选择)、谱变化稳健性技术、大种群性能、与其他说话人建模技术(单峰高斯、VQ码本、约束高斯混合和径向基函数)相比,混合高斯说话人模型在使用5秒纯净语音时获得了96.8%的识别准确率,在具有49个说话人群体的15秒电话语音中获得了80.8%的准确率,并且在相同的16个说话人电话语音任务上表现出优于其他说话人建模技术。
This paper introduces and motivates the use of Gaussian mixture models (GMM) for robust text-independent speaker identification, The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity, The focus of this work is on applications which require high identification rates using short utterance from unconstrained conversational speech and robustness to degradations produced by transmission over a telephone channel, A complete experimental evaluation of the Gaussian mixture speaker model is conducted on a 49 speaker, conversational telephone speech database, The experiments examine algorithmic issues (initialization, variance limiting, model order selection), spectral variability robustness techniques, large population performance, and comparisons to other speaker modeling techniques (uni-modal Gaussian, VQ codebook, tied Gaussian mixture, and radial basis functions), The Gaussian mixture speaker model attains 96.8% identification accuracy using 5 second clean speech utterances and 80.8% accuracy using 15 second telephone speech utterances with a 49 speaker population and is shown to outperform the other speaker modeling techniques on an identical 16 speaker telephone speech task.