Speaker Identification Within Whispered Speech Audio Streams

Speaker Identification Within Whispered Speech Audio Streams
复制标题

DOI:
10.1109/tasl.2010.2091631
复制
发表时间:
2011-07-01
影响因子:
--
通讯作者:
Hansen, John H. L.
Hansen, John H. L.
中科院分区:
其他
文献类型:
--
作者:
Fan, Xing;Hansen, John H. L.

文献摘要

被引文献

相似文献

耳语是自然会话中主体为保护隐私而使用的一种替代性语音产生模式。由于耳语和中性语音在激励和声道功能上的巨大差异,使用中性语音训练的说话人识别系统的性能显著下降。本文提出了一种无缝中性/耳语失配闭集说话人识别系统。首先,性能特性的中性训练闭集说话人识别系统的基础上梅尔频率倒谱系数高斯混合模型(MFCC GMM)的框架被认为是。据观察,耳语说话人识别,性能下降集中在只有一个子集的扬声器。接下来,它示出了在中性/耳语失配条件下的扬声器识别的性能损失集中在音素比低能量清音辅音。为了提高系统对清音辅音的识别性能,提出了一种基于线性和指数频率尺度的交替特征提取算法。分析了误识和正确识辨耳语的声学特性,以制定更有效的处理方案。提出了一个二维特征空间,以预测哪些耳语话语的系统将执行不佳,进行评估,以衡量耳语语音的质量。最后,提出了一个系统的无缝中性/耳语说话人识别,导致在8.85%-10.30%的绝对改善的说话人识别,与最好的封闭集的说话人ID性能的88.35%,获得了共961读耳语测试话语,和83.84%,共495自发耳语测试话语。
Whisper is an alternative speech production mode used by subjects in natural conversation to protect the privacy. Due to the profound differences between whisper and neutral speech in both excitation and vocal tract function, the performance of speaker identification systems trained with neutral speech degrades significantly. In this paper, a seamless neutral/whisper mismatched closed-set speaker recognition system is developed. First, performance characteristics of a neutral trained closed-set speaker ID system based on an Mel-frequency cepstral coefficient-Gaussian mixture model (MFCC-GMM) framework is considered. It is observed that for whisper speaker recognition, performance degradation is concentrated for only a subset of speakers. Next, it is shown that the performance loss for speaker identification in neutral/whisper mismatched conditions is focused on phonemes other than low-energy unvoiced consonants. In order to increase system performance for unvoiced consonants, an alternative feature extraction algorithm based on linear and exponential frequency scales is applied. The acoustic properties of misrecognized and correctly recognized whisper are analyzed in order to develop more effective processing schemes. A two-dimensional feature space is proposed in order to predict on which whispered utterances the system will perform poorly, with evaluations conducted to measure the quality of whispered speech. Finally, a system for seamless neutral/whisper speaker identification is proposed, resulting in an absolute improvement of 8.85%-10.30% for speaker recognition, with the best closed set speaker ID performance of 88.35% obtained for a total of 961 read whisper test utterances, and 83.84% using a total of 495 spontaneous whisper test utterances.