A prosody-based approach to end-of-utterance detection that does not require speech recognition

A prosody-based approach to end-of-utterance detection that does not require speech recognition
复制标题

一种基于韵律的语音结尾检测方法,不需要语音识别

DOI:
--
复制
发表时间:
2003
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
A. Stolcke
A. Stolcke
中科院分区:
--
文献类型:
--
作者:
Luciana Ferrer;Elizabeth Shriberg;A. Stolcke

文献摘要

被引文献

相似文献

在以前的工作中,我们表明,国家的最先进的话语结束检测(如使用,例如,在对话系统中)可以显着改善,通过使用韵律和/或语言模型,预测话语端点,基于单词和对齐输出从语音识别器。但是,在端点确定中使用识别器在某些应用中可能不实用。我们证明,由于韵律知识的改进,可以实现很大程度上没有对齐信息,即,而不需要语音识别器。一个韵律结束的话语检测器,只使用语音/非语音检测输出仍然是相当准确的,并具有较低的延迟比基线系统的基础上暂停长度阈值。
In previous work we showed that state-of-the-art end-of-utterance detection (as used, for example, in dialog systems) can be improved significantly by making use of prosodic and/or language models that predict utterance endpoints, based on word and alignment output from a speech recognizer. However, using a recognizer in endpointing might not be practical in certain applications. We demonstrate that the improvements due to the prosodic knowledge can be realized largely without alignment information, i.e., without requiring a speech recognizer. A prosodic end-of-utterance detector using only speech/nonspeech detection output is still considerably more accurate and has lower latency than a baseline system based on pause-length thresholding.