Canonicalization of feature parameters for automatic speech recognition

Canonicalization of feature parameters for automatic speech recognition
复制标题

自动语音识别的特征参数规范化

DOI:
10.21437/interspeech.2004-688
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
T. Nitta
T. Nitta
中科院分区:
--
文献类型:
--
作者:
Takashi Fukuda;T. Nitta

文献摘要

被引文献

相似文献

基于隐马尔可夫模型的分类器的声学模型包括各种类型的隐藏变量,如性别类型、语速和声学环境。如果存在一个规范化过程来减少来自AM的隐藏变量的影响,则可以实现稳健的自动语音识别(ASR)系统。在本文中,我们描述了以性别类型为隐藏变量的规范化过程的配置。该正则化过程由隐变量对应的多个不同语音特征(DPF)抽取器和DPF选择器组成,DPF选择器比较输入DPF和AM之间的距离。在DPF提取阶段,通过使用三个多层神经网络(MLN)将声学特征向量的输入序列映射到对应于男性、女性和中性声音的三个DPF空间。实验是通过比较(A)规范化DPF和单个HMM分类器的组合,以及(B)单个声学特征(MFCC)和多个HMM分类器的组合来进行的。结果表明,与传统的基于MFCC和单个隐马尔可夫模型的ASR以及基于多个隐马尔可夫模型的ASR相比,该方法具有更少的存储空间和更少的计算时间。
Acoustic models (AMs) of an HMM-based classifier include various types of hidden variables such as gender type, speaking rate, and acoustic environment. If there exists a canonicalization process that reduces the influence of the hidden variables from the AMs, a robust automatic speech recognition (ASR) system can be realized. In this paper, we describe the configuration of a canonicalization process targeting gender type as a hidden variable. The proposed canonicalization process is composed of multiple distinctive phonetic feature (DPF) extractors corresponding to the hidden variable and a DPF selector in which the distance between input DPF and AMs is compared. In a DPF extraction stage, an input sequence of acoustic feature vectors is mapped onto three DPF spaces corresponding to male, female, and neutral voice by using three multilayer neural networks (MLNs). Experiments are carried out by comparing (A) the combination of the canonicalized DPF and a single HMM classifier, and (B) the combination of a single acoustic feature (MFCC) and multiple HMM classifiers. The result shows that the proposed canonicalization method outperforms both of the conventional ASR with MFCC and a single HMM and the ASR with multiple HMMs in spite of less memories and computation time.