Enriching speech recognition with automatic detection of sentence boundaries and disfluencies

Enriching speech recognition with automatic detection of sentence boundaries and disfluencies
复制标题

DOI:
10.1109/tasl.2006.878255
复制
发表时间:
2006-09
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Yang Liu;Elizabeth Shriberg;A. Stolcke;D. Hillard;Mari Ostendorf;M. Harper
Yang Liu;Elizabeth Shriberg;A. Stolcke;D. Hillard;Mari Ostendorf;M. Harper
中科院分区:
其他
文献类型:
--
作者:
Yang Liu;Elizabeth Shriberg;A. Stolcke;D. Hillard;Mari Ostendorf;M. Harper

文献摘要

被引文献

相似文献

有效的人类和自动语音处理需要的不仅仅是单词的恢复。它还涉及恢复句子边界、填充词和不流畅等现象,这些现象称为结构元数据。我们描述了一个元数据检测系统,它结合了来自不同类型文本知识源的信息和来自韵律分类器的信息。我们研究了最大熵和条件随机场模型,以及主要的隐马尔可夫模型(HMM)方法,发现区分模型通常优于生成模型。我们报告了广播新闻和对话电话语音任务的系统性能,说明了任务之间的显著性能差异以及作为识别器性能的函数。这些结果代表了NIST RT-04F评估中所评估的最新技术
Effective human and automatic processing of speech requires recovery of more than just the words. It also involves recovering phenomena such as sentence boundaries, filler words, and disfluencies, referred to as structural metadata. We describe a metadata detection system that combines information from different types of textual knowledge sources with information from a prosodic classifier. We investigate maximum entropy and conditional random field models, as well as the predominant hidden Markov model (HMM) approach, and find that discriminative models generally outperform generative models. We report system performance on both broadcast news and conversational telephone speech tasks, illustrating significant performance differences across tasks and as a function of recognizer performance. The results represent the state of the art, as assessed in the NIST RT-04F evaluation