Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models

Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models
复制标题

通过混合方言语言模型进行混合方言的自动语音识别

DOI:
10.1109/taslp.2014.2387414
复制
发表时间:
2015
期刊:
EEE/ACM Transactions on Audio, Speech and Language Processing
影响因子:
--
通讯作者:
and Hiroshi G. Okuno
and Hiroshi G. Okuno
中科院分区:
--
文献类型:
--
作者:
Naoki Hirayama;Koichiro Yoshino;Katsutoshi Itoyama;Shinsuke Mori;and Hiroshi G. Okuno

文献摘要

相似文献

本文提出了一种可接受多种方言的自动语音识别系统。该系统通过对词汇变换的统计模拟和多种方言模型的组合来识别方言话语。以前的方言ASR系统是基于手工制作的几种方言词典,这涉及到昂贵的过程。该系统统计训练了通用语言和方言之间的转换规则,并基于机器翻译技术模拟了用于自动语音识别的方言语料库。这些规则是用小的平行语料库集来训练的,以弥补方言语言资源的不足。所提出的系统也接受包含各种词汇的混合方言话语。事实上,口语不是单一的方言,而是受说话者背景环境(如父母的母语或居住地)影响的混合方言。我们提出了两种方法,以适当地结合几种方言的每个发言者。第一种方法是使用混合方言的语言模型进行识别,并自动估计权重,使识别可能性最大化。该方法表现最好,但计算成本非常高,因为它对方言混合比例组合进行网格搜索,使识别可能性最大化。第二步是对各个单一方言语言模型的识别结果进行整合。与第一种方法相比,该模型的改进略小。然而,它的计算成本不高,并且可以在一般工作站上实时工作。两种方法对所有说话人的识别精度都高于单一方言模型和通用语言模型,我们可以在考虑计算成本和识别精度的情况下选择合适的模型用于ASR。
This paper presents an automatic speech recognition (ASR) system that accepts a mixture of various kinds of dialects. The system recognizes dialect utterances on the basis of the statistical simulation of vocabulary transformation and combinations of several dialect models. Previous dialect ASR systems were based on handcrafted dictionaries for several dialects, which involved costly processes. The proposed system statistically trains transformation rules between a common language and dialects, and simulates a dialect corpus for ASR on the basis of a machine translation technique. The rules are trained with small sets of parallel corpora to make up for the lack of linguistic resources on dialects. The proposed system also accepts mixed dialect utterances that contain a variety of vocabularies. In fact, spoken language is not a single dialect but a mixed dialect that is affected by the circumstances of speakers’ backgrounds (e.g., native dialects of their parents or where they live). We addressed two methods to combine several dialects appropriately for each speaker. The first was recognition with language models of mixed dialects with automatically estimated weights that maximized the recognition likelihood. This method performed the best, but calculation was very expensive because it conducted grid searches of combinations of dialect mixing proportions that maximized the recognition likelihood. The second was integration of results of recognition from each single dialect language model. The improvements with this model were slightly smaller than those with the first method. Its calculation cost was, however, inexpensive and it worked in real-time on general workstations. Both methods achieved higher recognition accuracies for all speakers than those with the single dialect models and the common language model, and we could choose a suitable model for use in ASR that took into consideration the computational costs and recognition accuracies.