Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model

Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model
复制标题

DOI:
10.1109/icassp.2018.8462576
复制
发表时间:
2018-04
期刊:
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Sei Ueno;H. Inaguma;M. Mimura;Tatsuya Kawahara
Sei Ueno;H. Inaguma;M. Mimura;Tatsuya Kawahara
中科院分区:
其他
文献类型:
--
作者:
Sei Ueno;H. Inaguma;M. Mimura;Tatsuya Kawahara

文献摘要

被引文献

相似文献

本文讨论了端到端的语音识别,直接映射到一个词序列的声学特征。声学到单词模型是有吸引力的,因为它不需要外部语言模型和复杂的解码器,从而导致非常简单和快速的解码。这种建模的明显缺点是训练数据的稀疏性,特别是对于不太频繁的单词。在本文中,我们提出了一个框架补充了字符级模型。单词级模型与字符级模型的联合训练增强了特征提取和分类过程的深度学习的通用性,防止其过拟合。此外,字符级模型被用来解码词汇表外(OOV)的单词,没有被词级模型覆盖。由于在端到端识别中有连接主义时间分类(CTC)和基于注意力的模型的选择,我们还探索了混合系统的最佳组合。对自发日语语料库(CSJ)的评估表明:(1)基于声音到单词注意力的模型优于CTC,(2)具有字符级CTC模型的多任务学习(MTL)是有效的,(3)混合系统实现了与标准DNN-HMM系统相当甚至更好的准确性,解码速度快25倍。
This paper addresses end-to-end speech recognition which directly maps acoustic features to a word sequence. The acoustic-to-word model is attractive since it does not require an external language model and an elaborate decoder, resulting in extremely simple and fast decoding. The apparent drawback of this modeling is sparseness of training data, particularly for less frequent words. In this paper, we propose a framework complemented with a character-level model. Joint training of the word-level model with the character-level model enhances the generality of deep learning of feature extraction and classification processes, preventing it from overfitting. Moreover, the character-level model is used to decode out-of-vocabulary (OOV) words that are not covered by the word-level model. Since there are choices of connectionist temporal classification (CTC) and attention-based models in the end-to-end recognition, we also explore optimal combination for the hybrid system. Evaluations on the Corpus of Spontaneous Japanese (CSJ) show that (1) the acoustic-to-word attention-based model outperforms CTC, (2) multitask learning (MTL) with character-level CTC model is effective, and (3) the hybrid system achieves comparable or even better accuracy than the standard DNN-HMM system with a decoding speed faster by a factor of 25.