Exploring the Use of an Unsupervised Autoregressive Model as a Shared Encoder for Text-Dependent Speaker Verification

Exploring the Use of an Unsupervised Autoregressive Model as a Shared Encoder for Text-Dependent Speaker Verification
复制标题

DOI:
10.21437/interspeech.2020-2957
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Vijay Ravi;Ruchao Fan;Amber Afshan;Huanhua Lu;A. Alwan
Vijay Ravi;Ruchao Fan;Amber Afshan;Huanhua Lu;A. Alwan
中科院分区:
其他
文献类型:
--
作者:
Vijay Ravi;Ruchao Fan;Amber Afshan;Huanhua Lu;A. Alwan

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的方法来解决文本相关的自动说话人确认(TD-ASV)通过使用一个共享的编码器与特定任务的解码器。自回归预测编码(APC)编码器以无监督的方式使用域外(LibriSpeech,VoxCeleb)和域内(DeepMine)未标记数据集进行预训练,以学习封装扬声器和语音内容的通用高级特征表示。使用标记数据集训练两个特定于任务的解码器来分类说话者(SID)和短语(PID)。使用PLDA对从SID解码器提取的说话者嵌入进行评分。SID和PID系统在评分水平上融合。与跨语言DeepMine数据集上的完全监督x向量基线相比,我们的系统的minDCF相对提高了51.9%。然而,i-vector/HMM方法优于所提出的APC编码器-解码器系统。在PID融合之前,x向量/PLDA基线和SID/PLDA分数的融合进一步将性能提高了15%,表明所提出的方法与x向量系统的互补性。我们表明,所提出的方法可以利用大型,未标记,数据丰富的域,并学习语音模式独立于下游任务。这样的系统可以在测试数据来自数据稀缺域的域不匹配场景中提供有竞争力的性能。
In this paper, we propose a novel way of addressing text-dependent automatic speaker verification (TD-ASV) by using a shared-encoder with task-specific decoders. An autoregressive predictive coding (APC) encoder is pre-trained in an unsupervised manner using both out-of-domain (LibriSpeech, VoxCeleb) and in-domain (DeepMine) unlabeled datasets to learn generic, high-level feature representation that encapsulates speaker and phonetic content. Two task-specific decoders were trained using labeled datasets to classify speakers (SID) and phrases (PID). Speaker embeddings extracted from the SID decoder were scored using a PLDA. SID and PID systems were fused at the score level. There is a 51.9% relative improvement in minDCF for our system compared to the fully supervised x-vector baseline on the cross-lingual DeepMine dataset. However, the i-vector/HMM method outperformed the proposed APC encoder-decoder system. A fusion of the x-vector/PLDA baseline and the SID/PLDA scores prior to PID fusion further improved performance by 15% indicating complementarity of the proposed approach to the x-vector system. We show that the proposed approach can leverage from large, unlabeled, data-rich domains, and learn speech patterns independent of downstream tasks. Such a system can provide competitive performance in domain-mismatched scenarios where test data is from data-scarce domains.