Dialogue act modeling for automatic tagging and recognition of conversational speech

Dialogue act modeling for automatic tagging and recognition of conversational speech
复制标题

DOI:
10.1162/089120100561737
复制
发表时间:
2000-09-01
影响因子:
9.3
通讯作者:
Meteer, M
Meteer, M
中科院分区:
计算机科学3区
文献类型:
--
作者:
Stolcke, A;Ries, K;Meteer, M

文献摘要

被引文献

相似文献

我们描述了一种用于对会话语音中的对话行为进行建模的统计方法,即,类似言语行为的单元,例如声明、问题、后台通道、同意、不同意和道歉。我们的模型检测和预测的基础上的词汇,搭配,韵律线索,以及对话行为序列的话语连贯性的对话行为。对话模型是基于治疗会话的话语结构作为一个隐藏的马尔可夫模型和个人的对话行为的观察,从模型的状态。通过对话行为n-gram对可能的对话行为序列的约束进行建模,将统计对话语法与单词n-gram、决策自由和神经网络相结合,对每个对话行为的特殊词汇nl td韵律表现进行建模。我们开发了一个概率集成的语音识别与对话建模,以提高语音识别和对话行为分类的准确性。使用来自自发的人对人电话语音的Switchboard语料库的1,155个保守的大型手动标记数据库来训练和评估模型。我们实现了良好的对话行为标记准确率(65%基于错误,自动识别的单词和韵律,71%基于单词成绩单,相比之下,机会基线准确率为35%,人类准确率为84%)和单词识别错误的小幅减少。
We describe a statistical approach for modeling dialogue acts in conversational speech, i.e., speech-act-like units such as STATEMENT, QUESTION, BACKCHANNEL, AGREEMENT, DISAGREEMENT, and APOLOGY. Our model detects and predicts dialogue acts based on lexical, collocational, and prosodic cues, as well as on the discourse coherence of the dialogue act sequence. The dialogue model is based on treating the discourse structure of a conversation as a hidden Markov model and the individual dialogue acts as observations emanating from the model states. Constraints on the likely sequence of dialogue acts are modeled via a dialogue act n-gram. The statistical dialogue grammar is combined with word N-grams, decision frees, and neural networks modeling the idiosyncratic lexical nl td prosodic manifestations of each dialogue act. We develop a probabilistic integration of speech recognition with dialogue modeling, to improve both speech recognition and dialogue act classification accuracy. Models are trained and evaluated using a large hand-labeled database of 1,155 conservations from the Switchboard corpus of spontaneous human-to-human telephone speech. We achieved good dialogue act labeling accuracy (65% based on errorful, automatically recognized words and prosody, and 71% based on word transcripts, compared to a chance baseline accuracy of 35% and human accuracy of 84%) and a small reduction in word recognition error.