REDUCED WORD ERROR RATES

REDUCED WORD ERROR RATES
复制标题

降低单词错误率

DOI:
10.21437/interspeech.2008-112
复制
发表时间:
1997
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
Jonathan G. Fiscus
Jonathan G. Fiscus
中科院分区:
--
文献类型:
--
作者:
Jonathan G. Fiscus

文献摘要

被引文献

相似文献

本文介绍了 NIST 开发的一个系统,当多个 ASR 系统的输出可用时,该系统可生成复合自动语音识别 (ASR) 系统输出,并且在许多情况下,复合 ASR 输出的错误率低于任何单个系统。该系统实施“投票”或重新评分过程来协调 ASR 系统输出中的差异。我们将该系统称为 NIST 识别器输出投票错误减少 (ROVER) 系统。随着额外的知识源添加到 ASR 系统(例如声学和语言模型),错误率通常会降低。本文描述了一种识别后过程,该过程将多个 ASR 系统生成的输出建模为独立的知识源,这些知识源可以组合并用于生成错误率较低的输出。为了实现这一点,多个 ASR 系统的输出通过动态编程@P) 对齐的迭代应用被组合成一个单一的、最小成本的字转换网络 (WTN)。通过自动重新评分或“投票”过程来搜索生成的网络,该过程选择得分最低的输出序列。
This paper describes a system developed at NIST to produce a composite Automatic Speech Recognition (ASR) system output when the outputs of multiple ASR systems are available, and for which, in many cases, the composite ASR output has lower error rate than any of the individual systems. The system implements a "voting" or rescoring process to reconcile differences in ASR system outputs. We refer to this system as the NIST Recognizer Output Voting Error Reduction (ROVER) system. As additional knowledge sources are added to an ASR system, (e.g., acoustic and language models), error rates are typically decrleased. This paper describes a post-recognition process which models the output generated by multiple ASR systems as independent knowledge sources that can be combined and used to generate an output with reduced error rate. To accomplish this, the outputs of multiple of ASR systems are combined into a single, minimal cost word transition network (WTN) via iterative applications of dynamic programming @P) alignments. The resulting network is searched by an automatic rescoring or "voting" process that selects an output sequence with the lowest score.