Multilingual Adaptation of RNN Based ASR Systems

Multilingual Adaptation of RNN Based ASR Systems
复制标题

基于 RNN 的 ASR 系统的多语言适应

DOI:
10.1109/icassp.2018.8461614
复制
发表时间:
2017
期刊:
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
A. Waibel
A. Waibel
中科院分区:
--
文献类型:
--
作者:
Markus Müller;Sebastian Stüker;A. Waibel

文献摘要

被引文献

相似文献

在这项工作中,我们专注于基于递归神经网络(RNN)的多语言系统,使用连接主义时间分类(CTC)损失函数进行训练。使用一套多语种的声学单元会带来困难。为了解决这个问题,我们提出了语言特征向量(LFVs)来训练语言自适应多语言系统。与说话人自适应不同,语言自适应不仅需要应用于特征层,还需要应用于网络的更深层。因此,在这项工作中,我们扩展了我们以前的方法,引入了一种新的技术,我们称之为“调制”。基于这种方法,我们使用LFVs调制RNN的隐藏层。我们评估了这种方法在充分和低资源条件下,以及字形和电话为基础的系统。通过使用调制,可以在不同条件下实现较低的错误率。
In this work, we focus on multilingual systems based on recurrent neural networks (RNNs), trained using the Connectionist Temporal Classification (CTC) loss function. Using a multilingual set of acoustic units poses difficulties. To address this issue, we proposed Language Feature Vectors (LFV s) to train language adaptive multilingual systems. Language adaptation, in contrast to speaker adaptation, needs to be applied not only on the feature level, but also to deeper layers of the network. In this work, we therefore extended our previous approach by introducing a novel technique which we call “modulation”. Based on this method, we modulated the hidden layers of RNNs using LFVs. We evaluated this approach in both full and low resource conditions, as well as for grapheme and phone based systems. Lower error rates throughout the different conditions could be achieved by the use of the modulation.