Combining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages

Combining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages
复制标题

结合串联和混合系统,以改进低资源语言的语音识别和关键字识别

DOI:
--
复制
发表时间:
2014
期刊:
Interspeech
影响因子:
--
通讯作者:
M. Gales
M. Gales
中科院分区:
--
文献类型:
--
作者:
S. Rath;K. Knill;A. Ragni;M. Gales

文献摘要

被引文献

相似文献

版权所有© 2014 ISCA.近年来,人们对低资源语言的自动语音识别(ASR)和关键词定位(KWS)系统产生了极大的兴趣。这一研究方向的驱动力之一是IARPA Babel项目。本文使用Babel项目发布的数据,研究了通过将两种形式的深度神经网络ASR系统(Tandem和Hybrid)结合起来,为ASR和KWS获得的性能增益。基线系统描述了五种选择期1语言:阿萨姆语、孟加拉语、海地克里奥尔语、老挝语和祖鲁语。所有ASR系统都具有共同的属性,例如深度神经网络配置,以及基于丰富语音问题和状态位置根节点的决策树。混合和串联系统的基准ASR和KWS性能进行了比较的“完整”,约80小时的培训数据,有限的,约10小时的培训数据,语言包。通过将这两个系统结合在一起,可以在所有配置中为KWS获得一致的性能增益。
Copyright © 2014 ISCA. In recent years there has been significant interest in Automatic Speech Recognition (ASR) and KeyWord Spotting (KWS) systems for low resource languages. One of the driving forces for this research direction is the IARPA Babel project. This paper examines the performance gains that can be obtained by combining two forms of deep neural network ASR systems, Tandem and Hybrid, for both ASR and KWS using data released under the Babel project. Baseline systems are described for the five option period 1 languages: Assamese; Bengali; Haitian Creole; Lao; and Zulu. All the ASR systems share common attributes, for example deep neural network configurations, and decision trees based on rich phonetic questions and state-position root nodes. The baseline ASR and KWS performance of Hybrid and Tandem systems are compared for both the "full", approximately 80 hours of training data, and limited, approximately 10 hours of training data, language packs. By combining the two systems together consistent performance gains can be obtained for KWS in all configurations.