A prosody-based approach to end-of-utterance detection that does not require speech recognition
A prosody-based approach to end-of-utterance detection that does not require speech recognition
复制标题
一种基于韵律的语音结尾检测方法,不需要语音识别
DOI:
--
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
A. Stolcke
中科院分区:
文献类型:
--
作者:
Luciana Ferrer;Elizabeth Shriberg;A. Stolcke
In previous work we showed that state-of-the-art end-of-utterance detection (as used, for example, in dialog systems) can be improved significantly by making use of prosodic and/or language models that predict utterance endpoints, based on word and alignment output from a speech recognizer. However, using a recognizer in endpointing might not be practical in certain applications. We demonstrate that the improvements due to the prosodic knowledge can be realized largely without alignment information, i.e., without requiring a speech recognizer. A prosodic end-of-utterance detector using only speech/nonspeech detection output is still considerably more accurate and has lower latency than a baseline system based on pause-length thresholding.