Coupling a Generative Model With a Discriminative Learning Framework for Speaker Verification

Coupling a Generative Model With a Discriminative Learning Framework for Speaker Verification
复制标题

DOI:
10.1109/taslp.2021.3129360
复制
发表时间:
2021-01
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Xugang Lu;Peng Shen;Yu-Yu Tsao-Yu;H. Kawai
Xugang Lu;Peng Shen;Yu-Yu Tsao-Yu;H. Kawai
中科院分区:
其他
文献类型:
--
作者:
Xugang Lu;Peng Shen;Yu-Yu Tsao-Yu;H. Kawai

文献摘要

相似文献

说话人确认(SV)的任务是判断说话人是目标说话人还是冒名顶替者。在大多数SV的研究中,对数似然比(LLR)得分估计的基础上的生成概率模型的说话人的功能,并与阈值进行比较,作出决定。然而,生成式模型通常关注个体特征分布,不具有区分性的特征选择能力,并且容易被多余的特征分散注意力。SV,作为一个假设检验,可以被公式化为一个二元判别任务,其中可以应用基于神经网络的判别学习。在判别式学习中,可以借助标签监督来去除讨厌的特征。然而,判别式学习更关注分类边界,并且容易过拟合到训练集,这可能导致测试集的泛化能力差。在本文中,我们提出了一个混合学习框架,即,将联合贝叶斯(JB)生成模型结构和参数与用于SV的神经判别学习框架耦合。在混合框架中,建立了一个两分支的连体神经网络与密集层耦合的因子化仿射变换中使用的JB模型。JB模型中的LLR得分估计是根据判别学习框架中的距离度量来制定的。通过用JB模型的生成学习的模型参数初始化两分支神经网络,我们进一步用成对样本训练模型参数作为二元判别任务。此外,直接评价指标(DEM)的基础上最小经验贝叶斯风险(EBR)的SV设计和集成作为一个目标函数的判别学习。我们在野生扬声器(SITW)和Voxceleb上进行了SV实验。实验结果表明,我们提出的模型提高了性能,与最先进的SV模型相比,有很大的利润。
The task of speaker verification (SV) is to decide whether an utterance is spoken by a target or an imposter speaker. In most studies of SV, a log-likelihood ratio (LLR) score is estimated based on a generative probability model on speaker features, and compared with a threshold for making a decision. However, the generative model usually focuses on individual feature distributions, does not have the discriminative feature selection ability, and is easy to be distracted by nuisance features. The SV, as a hypothesis test, could be formulated as a binary discrimination task where neural network based discriminative learning could be applied. In discriminative learning, the nuisance features could be removed with the help of label supervision. However, discriminative learning pays more attention to classification boundaries, and is prone to overfitting to a training set which may result in bad generalization on a test set. In this paper, we propose a hybrid learning framework, i.e., coupling a joint Bayesian (JB) generative model structure and parameters with a neural discriminative learning framework for SV. In the hybrid framework, a two-branch Siamese neural network is built with dense layers that are coupled with factorized affine transforms as used in the JB model. The LLR score estimation in the JB model is formulated according to the distance metric in the discriminative learning framework. By initializing the two-branch neural network with the generatively learned model parameters of the JB model, we further train the model parameters with the pairwise samples as a binary discrimination task. Moreover, a direct evaluation metric (DEM) in SV based on minimum empirical Bayes risk (EBR) is designed and integrated as an objective function in the discriminative learning. We carried out SV experiments on Speakers in the wild (SITW) and Voxceleb. Experimental results showed that our proposed model improved the performance with a large margin compared with state of the art models for SV.