Discriminative Neural Embedding Learning for Short-Duration Text-Independent Speaker Verification

Discriminative Neural Embedding Learning for Short-Duration Text-Independent Speaker Verification
复制标题

用于短时文本独立说话人验证的判别式神经嵌入学习

DOI:
10.1109/taslp.2019.2928128
复制
发表时间:
2019
影响因子:
5.4
通讯作者:
Yu Kai
Yu Kai
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wang Shuai;Huang Zili;Qian Yanmin;Yu Kai

文献摘要

被引文献

相似文献

短时间与文本无关的说话人确认仍然是近年来的研究热点,基于深度神经网络的嵌入在这种情况下表现出了令人印象深刻的结果。良好的说话人嵌入要求同时具有较小的类内变化和较大的类间差异,这对说话人的辨别和泛化能力至关重要。现有的嵌入学习策略可以分为两个框架:多阶段的级联嵌入学习和直接利用谱特征的嵌入学习。我们提出了新的方法来实现更具区分性的说话人嵌入。在级联框架内,提出了一种基于神经网络的深度判别分析(DDA)来将I向量投影到更具区分性的嵌入。在直接嵌入框架中,使用了具有更高级中心损失和A-Softmax损失的深度模型,并在该框架中研究了焦散。最后,将传统的I-向量法和神经嵌入法与基于神经网络的DDA算法相结合,进一步提高了算法的性能。主要实验是在从SRE语料库生成的与文本无关的短时间说话人确认数据集上进行的。实验结果表明,该方法在短时文本无关说话人确认中具有较好的应用前景,且优于传统的I-向量基准线和神经嵌入基准线。与I向量基线相比,最佳嵌入实现了大约30%的相对EER降低,当与I向量系统结合时,这一点可以进一步增强。
Short duration text-independent speaker verification remains a hot research topic in recent years, and deep neural network based embeddings have shown impressive results in such conditions. Good speaker embeddings require the property of both small intra-class variation and large inter-class difference, which is critical for the ability of discrimination and generalization. Current embedding learning strategies can be grouped into two frameworks: “Cascade embedding learning” with multiple stages and “direct embedding learning” from spectral feature directly. We propose new approaches to achieve more discriminant speaker embeddings. Within the cascade framework, a neural network based deep discriminant analysis (DDA) is proposed to project i-vector to more discriminative embeddings. Within the direct embedding framework, a deep model with more advanced center loss and A-softmax loss is used, the focal loss is also investigated in this framework. Moreover, the traditional i-vector and neural embeddings are finally combined with neural network based DDA to achieve further gain. Main experiments are carried out on a short-duration text-independent speaker verification dataset generated from the SRE corpus. The results show that the newly proposed method is promising for short-duration text-independent speaker verification, and it is consistently better than traditional i-vector and neural embedding baselines. The best embeddings achieve roughly 30% relative EER reduction compared to the i-vector baseline, which could be further enhanced when combined with the i-vector system.