Finding consensus in speech recognition

Finding consensus in speech recognition
复制标题

DOI:
--
复制
发表时间:
2000
期刊:
2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No.03EX721)
影响因子:
--
通讯作者:
Eric Brill;L. Mangu
Eric Brill;L. Mangu
中科院分区:
其他
文献类型:
--
作者:
Eric Brill;L. Mangu

文献摘要

被引文献

相似文献

本论文探索了利用语音识别系统产生的词格中存在的信息来提高识别输出的准确性并获得一组替代假设的更清晰的表示的新方法。我们改变了标准的问题制定搜索一个大的句子假设的一个小的候选字集的本地搜索。我们的方法取代了单词级后验概率作为语音识别的目标函数,对应于常用的基于单词的错误度量。该方法的核心是一个聚类过程,确定相互支持和竞争的词假设在一个格子,构建一个总的顺序在所有的词假设。再加上从识别器分数计算的单词后验概率,这允许有效地提取期望最小化单词错误率的假设。因此,我们的方法克服了基于单词的性能指标和标准的MAP评分范式,这是基于的,这可能会导致次优的识别结果之间的不匹配。我们还表明,我们的方法可以作为一个有效的晶格压缩技术。它的成功来自于丢弃后验概率较低的链接并重组剩余链接以创建一组新假设的能力。Switchboard语料库和广播新闻的实验表明,这种方法的结果显着的字错误率降低,无论是在标准的MAP方法相比,以前的字错误最小化技术的基础上N-最好的列表。我们还报告显着减少晶格尺寸相比,传统使用的技术。从本质上讲,我们的方法是一个词后验概率的估计器,因此可以受益于一些其他任务,如单词定位和置信度注释。
This thesis explores new ways of utilizing the information existing in word lattices produced by speech recognition systems to improve the accuracy of the recognition output and obtain a more perspicuous representation of a set of alternative hypotheses. We change the standard problem formulation of searching among a large set of sentence hypotheses to a local search in a small set of word candidates. Our approach replaces sentence-level posterior probabilities with word-level posteriors as the objective function for speech recognition, corresponding to the word-based error metric commonly used. The core of the method is a clustering procedure that identifies mutually supporting and competing word hypotheses in a lattice, constructing a total order over all word hypotheses. Together with word posterior probabilities computed from recognizer scores, this allows an efficient extraction of the hypothesis that is expected to minimize the word error rate. Our approach thus overcomes the mismatch between the word-based performance metric and the standard MAP scoring paradigm which is sentence-based, that can lead to sub-optimal recognition results. We also show that our method can be used as an efficient lattice compression technique. Its success comes from the ability to discard links with low a posteriori probability and recombine the remaining ones to create a new set of hypotheses. Experiments on the Switchboard corpus and Broadcast News show that this approach results in significant word error rate reductions, both over the standard MAP approach and compared to a previous word error minimization technique based on N-best lists. We also report significant decrease in lattice size when compared with the conventionally used technique. In essence, our method is an estimator of word posterior probabilities, and as such could benefit a number of other tasks like word spotting and confidence annotation.