End-to-End Language Diarization for Bilingual Code-Switching Speech

End-to-End Language Diarization for Bilingual Code-Switching Speech
复制标题

双语语码转换语音的端到端语言二化

DOI:
--
复制
发表时间:
2021
期刊:
Interspeech
影响因子:
--
通讯作者:
S. Styles
S. Styles
中科院分区:
--
文献类型:
--
作者:
Hexin Liu;Leibny Paola García Perera;Xinyi Zhang;J. Dauwels;Andy W. H. Khong;S. Khudanpur;S. Styles

文献摘要

被引文献

相似文献

针对双语语码转换语音,我们提出了两种端到端神经fi算法来实现语言二元化。fiRST是一种BLSTM-E2E体系结构,它包括一组堆叠的双向LSTM来计算嵌入,并结合了深度聚类损失来强制对属于同一类的语言进行分组。第二种是XSA-E2E架构,它基于一个x向量模型,然后是一个自我注意编码器。前者将框架级特征编码为段级嵌入,而后者考虑所有这些嵌入以生成段级语言标签序列。我们在WSTCSMC 2020中共享任务B获得的数据集和我们从SEAM数据集中手工制作的模拟数据上对所提出的方法进行了评估。实验结果表明,与WSTCSMC 2020数据集中的基准算法相比,我们提出的XSA-E2E结构在等错误率和准确率方面都有12.1%的相对提高和7.4%的相对提高。我们提出的XSA-E2E架构在SEAM数据集的模拟数据上获得了89.84%的准确率和85.60%的基线。
We propose two end-to-end neural configurations for language diarization on bilingual code-switching speech. The first, a BLSTM-E2E architecture, includes a set of stacked bidirectional LSTMs to compute embeddings and incorporates the deep clustering loss to enforce grouping of languages belonging to the same class. The second, an XSA-E2E architecture, is based on an x-vector model followed by a self-attention encoder. The former encodes frame-level features into segment-level embeddings while the latter considers all those embed-dings to generate a sequence of segment-level language labels. We evaluated the proposed methods on the dataset obtained from the shared task B in WSTCSMC 2020 and our handcrafted simulated data from the SEAME dataset. Experimental results show that our proposed XSA-E2E architecture achieved a relative improvement of 12.1% in equal error rate and a 7.4% relative improvement on accuracy compared with the baseline algo-rithm in the WSTCSMC 2020 dataset. Our proposed XSA-E2E architecture achieved an accuracy of 89.84% with a baseline of 85.60% on the simulated data derived from the SEAME dataset.