Combining Articulatory Features with End-to-End Learning in Speech Recognition
Combining Articulatory Features with End-to-End Learning in Speech Recognition
复制标题
将发音特征与语音识别中的端到端学习相结合
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
S. Wermter
中科院分区:
文献类型:
--
作者:
Leyuan Qu;C. Weber;Egor Lakomkin;Johannes Twiefel;S. Wermter
End-to-end neural networks have shown promising results on large vocabulary continuous speech recognition (LVCSR) systems. However, it is challenging to integrate domain knowledge into such systems. Specifically, articulatory features (AFs) which are inspired by the human speech production mechanism can help in speech recognition. This paper presents two approaches to incorporate domain knowledge into end-to-end training: (a) fine-tuning networks which reuse hidden layer representations of AF extractors as input for ASR tasks; (b) progressive networks which combine articulatory knowledge by lateral connections from AF extractors. We evaluate the proposed approaches on the speech Wall Street Journal corpus and test on the eval92 standard evaluation dataset. Results show that both fine-tuning and progressive networks can integrate articulatory information into end-to-end learning and outperform previous systems.