Compositional Embedding Models for Speaker Identification and Diarization with Simultaneous Speech From 2+ Speakers

Compositional Embedding Models for Speaker Identification and Diarization with Simultaneous Speech From 2+ Speakers
复制标题

DOI:
10.1109/icassp39728.2021.9413752
复制
发表时间:
2020-10
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Zeqian Li;J. Whitehill
Zeqian Li;J. Whitehill
中科院分区:
其他
文献类型:
--
作者:
Zeqian Li;J. Whitehill

文献摘要

相似文献

我们提出了一种新的说话人划分方法,可以处理2人以上的重叠语音。我们的方法基于组合嵌入[1]:与标准的说话者嵌入方法(如x-vector[2])一样,组合嵌入模型包含一个函数f,用于将不同说话者的语音分离开来。此外,它们还包括一个组合函数g,用于计算嵌入空间中的集合并运算,从而推断输入音频中的扬声器集合。在使用合成的LibriSpeech数据进行多人说话人识别的实验中,该方法优于传统的只训练分离单个说话人(而不是说话人集)的嵌入方法。在AMI耳机混合语料库上的扬声器diarization实验中,我们达到了最先进的准确率(DER=22.93%),略好于之前的最佳结果([3]的23.82%)。
We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding models contain a function f that separates speech from different speakers. In addition, they include a composition function g to compute set-union operations in the embedding space so as to infer the set of speakers within the input audio. In an experiment on multi-person speaker identification using synthesized LibriSpeech data, the proposed method outperforms traditional embedding methods that are only trained to separate single speakers (not speaker sets). In a speaker diarization experiment on the AMI Headset Mix corpus, we achieve state-of-the-art accuracy (DER=22.93%), slightly better than the previous best result (23.82% from [3]).