SdSV Challenge 2020: Large-Scale Evaluation of Short-Duration Speaker Verification

SdSV Challenge 2020: Large-Scale Evaluation of Short-Duration Speaker Verification
复制标题

SdSV Challenge 2020:短期说话人验证的大规模评估

DOI:
--
复制
发表时间:
2020
期刊:
Interspeech
影响因子:
--
通讯作者:
L. Burget
L. Burget
中科院分区:
--
文献类型:
--
作者:
Hossein Zeinali;Kong Aik LEE;Jahangir Alam;L. Burget

文献摘要

被引文献

相似文献

现代的说话人验证方法将语音表示为fi横长嵌入(fi-x-Long Embedding)。在这些方法中,我们隐含地假设说话人的特征与口语内容无关。当给出suffi很长的话语时,这样的假设通常成立。在这种情况下,说话人嵌入,如i向量和x向量,已被证明是非常有效的。对于短持续时间(几秒左右)的语音话语,说话人嵌入已经表现出对语音内容的显著fiCant依赖。在这方面,组织了2020年Sdsv挑战赛,广泛关注系统基准和分析短时长说话人验证(Sdsv)(Sdsv)的不同程度的语音变异。除了依赖文本和独立于文本的任务外,这项挑战还包括一项不寻常的、不同于以往的跨语言说话者验证(fifi)任务(英语与波斯语)。本文描述了数据集和任务、评估规则和协议、性能指标、基线系统和挑战结果。我们还介绍了从评估中获得的见解和未来的研究方向。
Modern approaches to speaker verification represent speech utterances as fixed-length embeddings. With these approaches, we implicitly assume that speaker characteristics are independent of the spoken content. Such an assumption generally holds when sufficiently long utterances are given. In this context, speaker embeddings, like i-vector and x-vector, have shown to be extremely effective. For speech utterances of short duration (in the order of a few seconds), speaker embeddings have shown significant dependency on the phonetic content. In this regard, the SdSV Challenge 2020 was organized with a broad focus on systematic benchmark and analysis on varying degrees of phonetic variability on short-duration speaker verification (SdSV). In addition to text-dependent and text-independent tasks, the challenge features an unusual and difficult task of cross-lingual speaker verification (English vs. Persian). This paper describes the dataset and tasks, the evaluation rules and protocols, the performance metric, baseline systems, and challenge results. We also present insights gained from the evaluation and future research directions.