Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates

Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates
复制标题

DOI:
10.1016/j.specom.2009.10.001
复制
发表时间:
2010-03-01
影响因子:
3.2
通讯作者:
Manning, Christopher D.
Manning, Christopher D.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Goldwater, Sharon;Jurafsky, Dan;Manning, Christopher D.

文献摘要

被引文献

相似文献

尽管进行了多年的语音识别研究,但人们对哪些单词容易被误识别以及原因知之甚少。之前的研究表明,不常见的单词、简短的单词以及非常响亮或快速的语音会导致错误增加,但许多其他推测的错误原因(例如,附近的不流利、转首词、语音邻域密度)从未经过仔细测试。说话者之间的错误率存在巨大差异的原因在很大程度上仍然是个谜。我们使用混合效应回归模型,通过分析两个最先进的识别器在会话语音上的错误来研究这些因素和其他因素。错误率较高的单词包括那些具有极端韵律特征的单词、那些出现在开头或作为话语标记的单词,以及双重易混淆的单词对:听觉上相似的单词也具有相似的语言模型概率。不连续中断点之前的单词(第一个重复标记和片段之前的单词)也具有较高的错误率。最后,即使在考虑了其他因素之后,说话者的差异也会导致错误率的巨大差异,这表明说话者错误率的差异并不能完全用单词选择、流畅性或韵律特征的差异来解释。我们还提出,双重易混淆对,而不是高邻域密度,可以更好地解释人类语音处理中的语音邻域错误。 (C) 2009 Elsevier B.V. 保留所有权利。
Despite years of speech recognition research, little is known about which words tend to be misrecognized and why. Previous work has shown that errors increase for infrequent words, short words, and very loud or fast speech, but many other presumed causes of error (e.g., nearby disfluencies, turn-initial words, phonetic neighborhood density) have never been carefully tested. The reasons for the huge differences found in error rates between speakers also remain largely mysterious.Using a mixed-effects regression model, we investigate these and other factors by analyzing the errors of two state-of-the-art recognizers on conversational speech. Words with higher error rates include those with extreme prosodic characteristics, those occurring turn-initially or as discourse markers, and doubly confusable pairs: acoustically similar words that also have similar language model probabilities. Words preceding disfluent interruption points (first repetition tokens and words before fragments) also have higher error rates. Finally, even after accounting for other factors, speaker differences cause enormous variance in error rates, suggesting that speaker error rate variance is not fully explained by differences in word choice, fluency, or prosodic characteristics. We also propose that doubly confusable pairs, rather than high neighborhood density, may better explain phonetic neighborhood errors in human speech processing. (C) 2009 Elsevier B.V. All rights reserved.