Structured discriminative models using deep neural-network features

Structured discriminative models using deep neural-network features
复制标题

DOI:
10.1109/asru.2015.7404789
复制
发表时间:
2015-10
期刊:
2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)
影响因子:
--
通讯作者:
R. V. Dalen;Jingzhou Yang;Haipeng Wang;A. Ragni;Chao Zhang-;M. Gales
R. V. Dalen;Jingzhou Yang;Haipeng Wang;A. Ragni;Chao Zhang-;M. Gales
中科院分区:
其他
文献类型:
--
作者:
R. V. Dalen;Jingzhou Yang;Haipeng Wang;A. Ragni;Chao Zhang-;M. Gales

文献摘要

相似文献

最先进的语音识别器采用各种配置的神经网络。一个标准的(混合)语音识别器计算一个时间框架和状态的可能性,只使用数千个可能的神经网络输出中的一个。然而,整个输出向量携带信息。在本文中,来自最先进的语音识别器的特征被收集到给定特定上下文的每个电话中,并输入到判别对数线性模型中。对数线性模型是用条件极大似然或大边际准则来训练的。一个关键因素是对数线性模型参数的先验。先验的均值被设置为达到原始系统性能的点。然后,对数线性模型在单个系统的最先进性能之上提供了额外的提高。
State-of-the-art speech recognisers employ neural networks in various configurations. A standard (hybrid) speech recogniser computes the likelihood for one time frame and state, using only one out of thousands of possible neural-network outputs. However, the whole output vector carries information. In this paper, features from state-of-the-art speech recognisers are collected per phone given a particular context, and input to a discriminative log-linear model. The log-linear model is trained with conditional maximum likelihood or a large-margin criterion. A key element is the prior on the parameters of the log-linear model. The mean of the prior is set to the point where the performance of the original systems is attained. The log-linear model then provides an additional increase over the state-of-the-art performance of the individual systems.