Robust Speaker Recognition in Noisy Conditions

Robust Speaker Recognition in Noisy Conditions
复制标题

DOI:
10.1109/tasl.2007.899278
复制
发表时间:
2007-07
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
J. Ming;Timothy J. Hazen;James R. Glass;D. Reynolds
J. Ming;Timothy J. Hazen;James R. Glass;D. Reynolds
中科院分区:
其他
文献类型:
--
作者:
J. Ming;Timothy J. Hazen;James R. Glass;D. Reynolds

文献摘要

被引文献

相似文献

本文在假设语音信号受到环境噪声干扰,但噪声特征未知的情况下,研究了噪声环境下的说话人识别与确认问题。这项研究的部分动机是说话人识别技术在手持设备或互联网上的潜在应用。虽然这些技术承诺提供额外的生物识别安全层来保护用户,但此类系统的实际实施面临许多挑战。其中之一是环境噪音。由于这种系统的移动性,噪声源可能是高度时变的,并且可能是未知的。这提高了在没有关于噪声的信息的情况下对噪声稳健性的要求。本文提出了一种将多条件模型训练和特征缺失理论相结合的方法,对具有未知时间-谱特征的噪声进行建模。利用有限噪声变化的模拟噪声数据进行多条件训练,提供对噪声的等价性补偿,并应用缺失特征理论通过忽略给定训练条件外的噪声变化来改进补偿,从而减少训练和测试的失配。本文的重点是与现实世界应用程序的新模型的实现有关的几个问题。这些措施包括生成多条件训练数据来对有噪语音进行建模,组合不同的训练数据来优化识别性能,以及降低模型的复杂性。新算法在两个数据库上进行了测试,分别使用模拟和真实噪声语音数据。第一个数据库是对TIMIT数据库的重新开发,方法是在存在各种噪声类型的情况下重新记录数据,用于测试说话人识别模型,重点是各种噪声。第二个数据库是在现实噪声条件下收集的手持设备数据库,用于进一步验证模型以用于真实世界的说话人验证。将新模型与基线系统进行比较,发现新模型实现了更低的错误率。
This paper investigates the problem of speaker identification and verification in noisy conditions, assuming that speech signals are corrupted by environmental noise, but knowledge about the noise characteristics is not available. This research is motivated in part by the potential application of speaker recognition technologies on handheld devices or the Internet. While the technologies promise an additional biometric layer of security to protect the user, the practical implementation of such systems faces many challenges. One of these is environmental noise. Due to the mobile nature of such systems, the noise sources can be highly time-varying and potentially unknown. This raises the requirement for noise robustness in the absence of information about the noise. This paper describes a method that combines multicondition model training and missing-feature theory to model noise with unknown temporal-spectral characteristics. Multicondition training is conducted using simulated noisy data with limited noise variation, providing a ldquocoarserdquo compensation for the noise, and missing-feature theory is applied to refine the compensation by ignoring noise variation outside the given training conditions, thereby reducing the training and testing mismatch. This paper is focused on several issues relating to the implementation of the new model for real-world applications. These include the generation of multicondition training data to model noisy speech, the combination of different training data to optimize the recognition performance, and the reduction of the model's complexity. The new algorithm was tested using two databases with simulated and realistic noisy speech data. The first database is a redevelopment of the TIMIT database by rerecording the data in the presence of various noise types, used to test the model for speaker identification with a focus on the varieties of noise. The second database is a handheld-device database collected in realistic noisy conditions, used to further validate the model for real-world speaker verification. The new model is compared to baseline systems and is found to achieve lower error rates.