Learning Speaker-Specific Characteristics With a Deep Neural Architecture

Learning Speaker-Specific Characteristics With a Deep Neural Architecture
复制标题

DOI:
10.1109/tnn.2011.2167240
复制
发表时间:
2011-11-01
影响因子:
--
通讯作者:
Salman, Ahmad
Salman, Ahmad
中科院分区:
其他
文献类型:
--
作者:
Chen, Ke;Salman, Ahmad

文献摘要

被引文献

相似文献

语音信号传达各种各样的混合信息,从语言到说话者特定的信息。然而,大多数声学表示都将各种不同类型的信息作为一个整体来描述,这可能会阻碍语音或说话人识别(SR)系统产生更好的性能。在本文中,我们提出了一种新的深度神经架构(DNA),特别是用于从梅尔频率倒谱系数中学习特定于说话者的特征,梅尔频率倒谱系数是一种常用于语音识别和SR的声学表示,它会导致特定于说话者的过完备表示。为了学习内在的特定于说话人的特征,我们提出了一个目标函数,该目标函数由说话人相似性/不相似性方面的对比损失和数据重建损失组成,用于正则化非说话人相关信息的干扰。此外,我们采用混合学习策略来学习深度神经网络的参数:即,局部但贪婪的逐层无监督预训练用于初始化,全局监督学习用于最终的判别目标。通过四个语言数据联盟(LDC)基准和两个非英语语料库,我们证明了我们的过完备表示在表征各种说话者时是鲁棒的,无论他们的话语是否用于训练我们的DNA,并且对文本和语言高度不敏感。广泛的比较研究表明,我们的方法产生最喜爱的结果,在说话人确认和分割。最后,我们讨论了几个问题,我们提出的方法。
Speech signals convey various yet mixed information ranging from linguistic to speaker-specific information. However, most of acoustic representations characterize all different kinds of information as whole, which could hinder either a speech or a speaker recognition (SR) system from producing a better performance. In this paper, we propose a novel deep neural architecture (DNA) especially for learning speaker-specific characteristics from mel-frequency cepstral coefficients, an acoustic representation commonly used in both speech recognition and SR, which results in a speaker-specific overcomplete representation. In order to learn intrinsic speaker-specific characteristics, we come up with an objective function consisting of contrastive losses in terms of speaker similarity/dissimilarity and data reconstruction losses used as regularization to normalize the interference of non-speaker-related information. Moreover, we employ a hybrid learning strategy for learning parameters of the deep neural networks: i.e., local yet greedy layerwise unsupervised pretraining for initialization and global supervised learning for the ultimate discriminative goal. With four Linguistic Data Consortium (LDC) benchmarks and two non-English corpora, we demonstrate that our overcomplete representation is robust in characterizing various speakers, no matter whether their utterances have been used in training our DNA, and highly insensitive to text and languages spoken. Extensive comparative studies suggest that our approach yields favorite results in speaker verification and segmentation. Finally, we discuss several issues concerning our proposed approach.