WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding

WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding
复制标题

WFST 支持 ASR 问题的解决方案:超越 HMM 解码

DOI:
--
复制
发表时间:
2012
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Ney
H. Ney
中科院分区:
--
文献类型:
--
作者:
Björn Hoffmeister;G. Heigold;David Rybach;R. Schlüter;H. Ney

文献摘要

被引文献

相似文献

在过去的十年中,加权有限状态转换器(WFST)已成为流行的语音识别。虽然其主要应用领域仍然是隐马尔可夫模型(HMM)解码,但WFST框架现在也被视为自动语音识别(ASR)中许多其他核心问题的解决方案。这些解决方案是鲜为人知的,这项工作的目的是在大词汇量的连续语音识别(LVCSR)除了HMM解码的WFST应用的概述:歧视性的声学模型训练,贝叶斯风险解码,和系统组合。WFST框架的应用程序有很大的实际影响:我们展示了框架如何帮助结构化问题,开发通用的解决方案,并将复杂的计算委托给WFST工具包。在本文中,我们回顾文献,讨论现有的方法,并提供新的见解WFST启用的解决方案。我们还提出了一种新的,纯粹的WFST为基础的算法计算确切的贝叶斯风险假设从一个格子的Levenshtein距离作为损失函数。我们提出了一个统一的框架中的问题和解决方案,并讨论了使用WFST的优点和局限性。我们不提供新的实验结果,但参考现有的文献。我们的工作有助于确定换能器框架在哪里以及如何有助于LVCSR问题的紧凑和通用的解决方案。
During the last decade, weighted finite-state transducers (WFSTs) have become popular in speech recognition. While their main field of application remains hidden Markov model (HMM) decoding, the WFST framework is now also seen as a brick in solutions to many other central problems in automatic speech recognition (ASR). These solutions are less known, and this work aims at giving an overview of the applications of WFSTs in large-vocabulary continuous speech recognition (LVCSR) besides HMM decoding: discriminative acoustic model training, Bayes risk decoding, and system combination. The application of the WFST framework has a big practical impact: we show how the framework helps to structure problems, to develop generic solutions, and to delegate complex computations to WFST toolkits. In this paper, we review the literature, discuss existing approaches, and provide new insights into WFST enabled solutions. We also present a novel, purely WFST-based algorithm for computing the exact Bayes risk hypothesis from a lattice with the Levenshtein distance as loss function. We present the problems and their solutions in a unified framework and discuss the advantages and limits of using WFSTs. We do not provide new experimental results, but refer to the existing literature. Our work helps to identify where and how the transducer framework can contribute to a compact and generic solution to LVCSR problems.