Untranscribed Web Audio for Low Resource Speech Recognition

Untranscribed Web Audio for Low Resource Speech Recognition
复制标题

用于低资源语音识别的未转录网络音频

DOI:
--
复制
发表时间:
2019
期刊:
Interspeech
影响因子:
--
通讯作者:
S. Renals
S. Renals
中科院分区:
--
文献类型:
--
作者:
Andrea Carmantini;P. Bell;S. Renals

文献摘要

被引文献

相似文献

语音识别模型非常容易受到训练数据和评估数据之间声学和语言领域的失配的影响。对于低资源语言,很难获得目标领域的转录语音,而非转录数据可以用最小的努力来收集。最近,一种将无格子最大互信息(LF-MMI)应用于未转录数据的方法被发现对于半监督训练是有效的。然而,较弱的初始模型和区域失配会导致半监督模型的高删失率。因此,我们提出了一种方法,依靠LF-MMI处理不确定性的能力,迫使基本模型过度生成可能的转录。在IARPA资料计划的数据上,我们的新的半监督方法比标准的半监督方法性能更好,在适应不匹配的带宽和区域时产生了显著的收益。
Speech recognition models are highly susceptible to mismatch in the acoustic and language domains between the training and the evaluation data. For low resource languages, it is difficult to obtain transcribed speech for target domains, while untranscribed data can be collected with minimal effort. Recently, a method applying lattice-free maximum mutual information (LF-MMI) to untranscribed data has been found to be effective for semi-supervised training. However, weaker initial models and domain mismatch can result in high deletion rates for the semi-supervised model. Therefore, we propose a method to force the base model to overgenerate possible transcriptions, relying on the ability of LF-MMI to deal with uncertainty. On data from the IARPA MATERIAL programme, our new semi-supervised method outperforms the standard semisupervised method, yielding significant gains when adapting for mismatched bandwidth and domain.