Localization based Sequential Grouping for Continuous Speech Separation

Localization based Sequential Grouping for Continuous Speech Separation
复制标题

DOI:
10.1109/icassp43922.2022.9746896
复制
发表时间:
2021-07
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Zhong-Qiu Wang;Deliang Wang
Zhong-Qiu Wang;Deliang Wang
中科院分区:
其他
文献类型:
--
作者:
Zhong-Qiu Wang;Deliang Wang

文献摘要

相似文献

本研究针对连续语音分离与说话人日志化,探讨强健的说话人定位方法,使用说话人方向将同一说话人的非连续片段分组。假设说话人不移动,并位于不同的方向,到达方向(DOA)的信息提供了一个准确的顺序分组和说话人日记的信息线索。我们的系统在以下意义上是块在线的。给定最多两个扬声器的帧块,我们应用两个扬声器分离模型来分离(和增强)扬声器,估计每个分离的扬声器的DOA,并基于DOA估计将分离结果分组到块。在LibriCSS语料库上的说话人日志化和说话人属性语音识别结果验证了该算法的有效性。
This study investigates robust speaker localization for continuous speech separation and speaker diarization, where we use speaker directions to group non-contiguous segments of the same speaker. Assuming that speakers do not move and are located in different directions, the direction of arrival (DOA) information provides an informative cue for accurate sequential grouping and speaker diarization. Our system is block-online in the following sense. Given a block of frames with at most two speakers, we apply a two-speaker separation model to separate (and enhance) the speakers, estimate the DOA of each separated speaker, and group the separation results across blocks based on the DOA estimates. Speaker diarization and speaker-attributed speech recognition results on the LibriCSS corpus demonstrate the effectiveness of the proposed algorithm.