Combining Articulatory Features with End-to-End Learning in Speech Recognition

Combining Articulatory Features with End-to-End Learning in Speech Recognition
复制标题

将发音特征与语音识别中的端到端学习相结合

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Artificial Neural Networks
影响因子:
--
通讯作者:
S. Wermter
S. Wermter
中科院分区:
--
文献类型:
--
作者:
Leyuan Qu;C. Weber;Egor Lakomkin;Johannes Twiefel;S. Wermter

文献摘要

被引文献

相似文献

端到端神经网络在大词汇量连续语音识别(LVCSR)系统中显示出了很好的效果。然而,它是具有挑战性的领域知识集成到这样的系统。具体而言,发音特征(AF)的灵感来自人类的语音产生机制,可以帮助语音识别。本文提出了两种将领域知识纳入端到端训练的方法:(a)微调网络,它重用AF提取器的隐藏层表示作为ASR任务的输入;(B)渐进网络,它通过AF提取器的横向连接将联合收割机发音知识结合起来。我们评估所提出的方法的语音华尔街日报语料库和eval 92标准评估数据集上进行测试。结果表明,微调和渐进式网络都可以将发音信息集成到端到端学习中,并且优于以前的系统。
End-to-end neural networks have shown promising results on large vocabulary continuous speech recognition (LVCSR) systems. However, it is challenging to integrate domain knowledge into such systems. Specifically, articulatory features (AFs) which are inspired by the human speech production mechanism can help in speech recognition. This paper presents two approaches to incorporate domain knowledge into end-to-end training: (a) fine-tuning networks which reuse hidden layer representations of AF extractors as input for ASR tasks; (b) progressive networks which combine articulatory knowledge by lateral connections from AF extractors. We evaluate the proposed approaches on the speech Wall Street Journal corpus and test on the eval92 standard evaluation dataset. Results show that both fine-tuning and progressive networks can integrate articulatory information into end-to-end learning and outperform previous systems.