The STC System for the CHiME-6 Challenge

The STC System for the CHiME-6 Challenge
复制标题

用于 CHiME-6 挑战的 STC 系统

DOI:
--
复制
发表时间:
2020
期刊:
6th International Workshop on Speech Processing in Everyday Environments (CHiME 2020)
影响因子:
--
通讯作者:
A. Romanenko
A. Romanenko
中科院分区:
--
文献类型:
--
作者:
I. Medennikov;M. Korenevsky;Tatiana Prisyach;Yuri Y. Khokhlov;Mariya Korenevskaya;Ivan Sorokin;Tatiana Timofeeva;Anton Mitrofanov;A. Andrusenko;Ivan Podluzhny;A. Laptev;A. Romanenko

文献摘要

被引文献

相似文献

本文描述了针对CHINE-6挑战的语音技术中心(STC)系统,该挑战旨在晚宴场景中的多麦克风、多说话人语音识别和二元化。我们参加了赛道1和赛道2,并提交了每个赛道的排名A和排名B的结果。一轨系统采用了基于软活动的导引源分离(GSS)作为前端,结合了基于GSS的训练数据增强、多步长多流自关注层、统计层和谱增强等先进的声学建模技术,以及声学模型的格子级融合。我们的Track 1系统排在前三位,在基线的基础上实现了30%的相对WER降低。此外,使用神经语言模型对B进行格点重新排序。总体来说,这导致轨道1的相对WER比基线降低了34%。对于轨道2,我们提出了一种新的目标-说话人语音活动检测(TS-VAD)方法来解决二值化问题。良好的二值化结果使得在获得的片段上执行GSS成为可能。TS-VAD基于i向量说话人嵌入,最初使用基于x向量谱聚类的强二值化系统进行估计。第二个轨道使用来自轨道1系统的后端。赛道2的系统表现出了最先进的性能,相对地超过了基线39%的DER、45%的JER、43%的WER(排名A)和45%的WER(排名B)。
This paper is a description of the Speech Technology Center (STC) systems for the CHiME-6 challenge aimed at multi-microphone multi-speaker speech recognition and diarization in a dinner party scenario. We participated in both Track 1 and Track 2 and submitted our results for Ranking A as well as Ranking B for each track. The soft-activity based Guided Source Separation (GSS) as a front-end and a combination of advanced acoustic modeling techniques such as GSS-based training data augmentation, multi-stride and multi-stream self-attention layers, statistics layer and SpecAugment, as well as the lattice-level fusion of acoustic models were applied in the 1st track system. Our system for Track 1 was in the top three systems, achieving 30% relative WER reduction over the baseline. Additionally, lattice rescoring with a neural language model was applied for Ranking B. Overall, this led to 34% relative WER reduction over the baseline in Track 1. For Track 2, we proposed a novel Target-Speaker Voice Activity Detection (TS-VAD) approach to solve the diarization problem. Good diarization results made it possible to perform GSS on the obtained segments. TS-VAD is based on i-vector speaker embeddings, which are initially estimated using a strong diarization system based on spectral clustering of x-vectors. The back-end from the Track 1 system was used in the second track. The system for Track 2 demonstrated state-of-the-art performance, outperforming the baseline by 39% DER, 45% JER, 43% WER (Ranking A) and 45% WER (Ranking B) relative.