Prosody-based detection of the context of backchannel responses

Prosody-based detection of the context of backchannel responses
复制标题

基于韵律的反向通道响应上下文检测

DOI:
10.21437/icslp.1998-71
复制
发表时间:
1998
期刊:
5th International Conference on Spoken Language Processing (ICSLP 1998)
影响因子:
--
通讯作者:
Yasuharu Den
Yasuharu Den
中科院分区:
--
文献类型:
--
作者:
Hiroaki Noguchi;Yasuharu Den

文献摘要

被引文献

相似文献

当前的口语对话系统缺乏正反馈,如反向渠道,这在人与人之间的对话中很常见。为了开发更自然的人机界面,对反向通道响应的研究是必不可少的。在本文中,我们提出了一种检测日语中反向通道响应精确计时的方法,旨在将这种方法纳入未来的口语对话系统中。该方法基于具有多种韵律特征的机器学习技术。结果表明,该方法可以有效地自动推导出检测回信道上下文的规则。我们的方法的性能比以前的方法好得多。1. 许多研究人员已经报道,由于缺乏系统的积极反馈,人们在与语音对话系统交谈时犹豫不决,如反向通道,这在人与人之间的对话中很常见[3,6]。为了开发更自然的人机界面,对反向通道响应机制的研究是必不可少的。在本文中,我们提出了一种检测日语反信道响应精确定时的方法,旨在将这种方法纳入未来的口语对话系统中。在该方法中,仅使用基本频率和能量等韵律特征来检测反向信道的上下文,这是当前语音技术相对容易处理的。与使用非常有限的特征和手工启发式的现有方法相比,我们采用了一种具有各种韵律特征的机器学习方法,这些特征可能与反向通道上下文的检测相关。将证明我们的方法在自动导出检测反向信道上下文的规则方面是有效的,并且它的性能比以前的方法要好得多。在第二节中,我们回顾了日语会话中的反向信道和反向信道时间自动检测的相关工作。在第3节中,我们描述了我们研究中使用的口语对话语料库,并提供了我们对反向通道的定义。在第4节中,我们进行了一个心理学实验,以便对普通人常见的反向渠道的积极和消极背景进行分类。在第5节中,我们使用决策树学习方法获得了最能区分反向通道的积极和消极语境的韵律线索。第六部分对全文进行了总结。
ABSTRACT Current spoken dialogue systems lack positive feedback such asbackchannels, which are common in human-human conversa-tions. To develop more natural human-computer interfaces, theinvestigation of backchannel-responses are indispensable. In thispaper, we propose a method for detecting the precise timing forbackchannel responses in Japanese and aim at incorporating suchmethod in future spoken dialogue systems. The proposed methodis based on machine learning technique with a variety of prosodicfeatures. It is shownto be effectivein automatically derivingrulesfor detecting the contexts of backchannels. The performance ofour method is considerably better than previous methods. 1. INTRODUCTION Many researchers have reported that people hesitate to talk withspokendialogue systems due to the lack of positivefeedback fromthe systems such as backchannels, which are common in human-human conversations [3, 6]. To develop more natural human-computer interfaces, the investigation of backchannel-responsemechanisms are indispensable. In this paper, we propose amethodfordetecting the precisetiming forbackchannel responsesin Japanese and aim at incorporating such method in future spo-ken dialogue systems.In the proposed method, the contexts for backchannels are de-tected by using only prosodic features such as fundamental fre-quency and energy, which are relatively easy to handle by currentspeech technology. In contrast to the existing methods, whichuse very limited number of features and hand-made heuristics, weemploy a machine learning method with a varietyof prosodic fea-tures which might be relevant to the detection of the backchannelcontext. It will be shown that our method is effective in automati-cally deriving rules for detecting the contextsof backchannels andthat it performs considerably better than previous methods.In Section 2, we review related works on backchannels inJapanese conversation and automatic detection of the timing forbackchannels. In Section 3, we describe the spoken dialogue cor-pus used in our study and provide our definition of backchannels.In Section 4, we conduct a psychological experiment in order tocategorize positive and negative contexts for backchannels whichare common to average humans. In Section 5, we obtain, by us-ing decision tree learning method, prosodic cues which best dis-criminate the positive and negative contexts for backchannels. InSection 6, we summarize the paper.