SURT 2.0: Advances in Transducer-Based Multi-Talker Speech Recognition

SURT 2.0: Advances in Transducer-Based Multi-Talker Speech Recognition
复制标题

DOI:
10.1109/taslp.2023.3318398
复制
发表时间:
2023-06
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Desh Raj;Daniel Povey;S. Khudanpur
Desh Raj;Daniel Povey;S. Khudanpur
中科院分区:
其他
文献类型:
--
作者:
Desh Raj;Daniel Povey;S. Khudanpur

文献摘要

相似文献

流分解和识别换能器(SURT)模型是最近提出的用于连续、流、多说话者语音识别(ASR)的端到端方法。尽管SURT在多轮会议上取得了令人印象深刻的结果,但它有明显的局限性:(I)它存在泄漏和遗漏相关的错误;(Ii)它的计算成本很高,因此没有在学术界得到采用;(Iii)它只在合成混合物上进行评估。在这项工作中,我们提出了对原始Surt的几个修改,这些修改是精心设计的,以解决上述限制。具体地说,我们(I)将解混模块改变为使用双路径建模的掩码估计器,(Ii)对换能器使用流Zipform编码器和无状态解码器,(Iii)使用力对齐子段执行混合模拟,(Iv)在单说话人数据上预训练换能器,(V)使用掩蔽损失和编码器CTC损失形式的辅助目标,以及(Vi)执行域自适应用于远场识别。我们表明,我们的修改允许Surt 2.0在多人ASR结果方面优于其前身,同时足够有效地利用学术资源进行训练。我们在3个公开可用的会议基准上进行了评估-LibriCS、AMI和ICSI,其中我们的最佳模型在远场不分段记录上的WER值分别为16.9%、44.6%和32.2%。
The Streaming Unmixing and Recognition Transducer (SURT) model was proposed recently as an end-to-end approach for continuous, streaming, multi-talker speech recognition (ASR). Despite impressive results on multi-turn meetings, SURT has notable limitations: (i) it suffers from leakage and omission related errors; (ii) it is computationally expensive, due to which it has not seen adoption in academia; and (iii) it has only been evaluated on synthetic mixtures. In this work, we propose several modifications to the original SURT which are carefully designed to fix the above limitations. In particular, we (i) change the unmixing module to a mask estimator that uses dual-path modeling, (ii) use a streaming zipformer encoder and a stateless decoder for the transducer, (iii) perform mixture simulation using force-aligned subsegments, (iv) pre-train the transducer on single-speaker data, (v) use auxiliary objectives in the form of masking loss and encoder CTC loss, and (vi) perform domain adaptation for far-field recognition. We show that our modifications allow SURT 2.0 to outperform its predecessor in terms of multi-talker ASR results, while being efficient enough to train with academic resources. We conduct our evaluations on 3 publicly available meeting benchmarks — LibriCSS, AMI, and ICSI, where our best model achieves WERs of 16.9%, 44.6% and 32.2%, respectively, on far-field unsegmented recordings.