Continuous Speech Separation with Recurrent Selective Attention Network

Continuous Speech Separation with Recurrent Selective Attention Network
复制标题

具有循环选择性注意网络的连续语音分离

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Jinyu Li
Jinyu Li
中科院分区:
--
文献类型:
--
作者:
Yixuan Zhang;Zhuo Chen;Jian Wu;Takuya Yoshioka;Peidong Wang;Zhong Meng;Jinyu Li

文献摘要

被引文献

相似文献

基于置换不变量训练(PIT)的连续语音分离(CSS)虽然显著提高了会话转录的准确率,但由于其输出通道数固定,在热点区域经常存在语音泄漏和分离失败的问题。在本文中,我们提出将递归选择性注意网络(RSAN)应用于说话人识别系统中,根据说话人的活动计数产生不同数目的输出通道。此外,通过在CSS框架中引入相邻处理块之间的依赖关系,提出了一种新的基于块的RSAN依赖扩展。它使网络能够利用来自先前块的分离结果来促进当前块处理。在LibriCSS数据集上的实验结果表明,基于RSAN的CSS(RSAN-CSS)网络与基于PIT的模型相比,语音识别的准确率得到了一致的提高。提出的分块依赖模型进一步提高了RSAN-CS的性能。
While permutation invariant training (PIT) based continuous speech separation (CSS) significantly improves the conversation transcription accuracy, it often suffers from speech leakages and failures in separation at "hot spot" regions because it has a fixed number of output channels. In this paper, we propose to apply recurrent selective attention network (RSAN) to CSS, which generates a variable number of output channels based on active speaker counting. In addition, we propose a novel block-wise dependency extension of RSAN by introducing dependencies between adjacent processing blocks in the CSS framework. It enables the network to utilize the separation results from the previous blocks to facilitate the current block processing. Experimental results on the LibriCSS dataset show that the RSAN-based CSS (RSAN-CSS) network consistently improves the speech recognition accuracy over PIT-based models. The proposed block-wise dependency modeling further boosts the performance of RSAN-CSS.