Incremental Dialogue System Faster than and Preferred to its Nonincremental Counterpart

Incremental Dialogue System Faster than and Preferred to its Nonincremental Counterpart
复制标题

增量对话系统比非增量对话系统更快且更受青睐

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
Micheal K. Tanenhaus
Micheal K. Tanenhaus
中科院分区:
--
文献类型:
--
作者:
Gregory Aist;James F. Allen;E. Campana;C. G. Gallo;S. Stoness;Mary D. Swift;Micheal K. Tanenhaus

文献摘要

被引文献

相似文献

增量对话系统比其非增量对话系统更快且更受欢迎,与其非增量对话系统相比,Gregory.aist 1(gregory.aist@asu.edu)、James Allen 2(James@cs.rochester.edu)、Ellen Campana 2、3、4、5(ecampana@bcs.rochester.edu)、Carlos Gomez Gallo 2(cgomez@cs.rochester.edu)、Scott Stoness 2(Stoness@cs.rochester.edu)、Mary Swift 2(wift@cs.rochester.edu)、和Michael K.Tanenhaus 3(mtan@bcs.rochester.edu)亚利桑那州立大学计算机科学与工程系,邮编:878809,坦佩,AZ,85287,罗切斯特大学计算机科学系,邮编:270226,纽约州罗切斯特,14627,罗切斯特大学,脑与认知科学系,邮编:270268,邮编:14627;艺术、媒体和工程项目亚利桑那州立大学邮政信箱878709坦佩亚利桑那州立大学85287心理学系亚利桑那州立大学邮政信箱871104坦佩亚利桑那州立大学85287根据听众的注视确定(阿尔特曼和凯姆1999年)。还可以基于部分话语来采取其他动作。有许多不同的知识来源可用于理解。在语音识别方面,常用的信息源包括声学、语音学和音位学、词汇概率和词序。在对话系统中,额外的信息源通常包括语法和语义(一般的和特定于领域的)。然而,也有一些信息来源不太频繁地编程。这些信息包括词法和韵律等语言信息。基于知识的功能也是可用的,比如世界知识(三角形有三条边)、领域知识(这里有两个大小的三角形)和任务知识(下一步是点击一个小三角形)。还可以从视觉环境中获得实用信息(旗帜附近有一个小三角形。)在本文中,我们讨论了我们在建立机器对口语的增量理解方法方面取得的一些进展。我们首先讨论我们和其他人在这一领域的一些相关工作。然后,我们讨论了我们一直在开发的试验床领域,并展示了该领域中人与人对话的一些特征。然后,我们讨论我们一直在开发的增量式体系结构,强调它与传统体系结构的区别。最后,我们给出了一个系统性能的实验评估,表明增量系统比非增量系统更快,也更受欢迎。当前的对话系统通常以流水线、模块化的方式一次对一个完整的发声进行操作。来自人类语言理解的证据表明,人类理解是循序渐进的,并在句法分析过程中利用多种信息来源,包括传统上的“后来”成分,如语用学。在本文中,我们描述了一个口语对话系统,它增量地理解语言,在用户说话过程中提供关于可能的指代物的可视反馈,并允许重叠的言语和动作。我们进一步介绍了一项实证研究的结果,表明由此产生的对话系统总体上比它的非增量对应系统更快。此外,增量式系统比非增量式系统更受欢迎--超出了速度和准确性等因素的影响。这些结果表明,成功的增量式理解系统将提高性能和可用性。关键词:自然语言系统;增量处理。对话导论对话系统自然语言理解的标准模型是流水线的、模块化的,并在完整的话语上运行。我们所说的流水线是指一次只有一个级别的处理以顺序方式运行。通过模块化,我们的意思是每个级别的处理仅依赖于前一个级别。我们所说的完整话语,是指系统一次只处理一句话。然而,有相当多的证据表明,人类的语言处理既不是流水线的,也不是模块化的,也不是完整的(Marslen-Wilson,1993)。证据从各种来源汇聚在一起,特别是在演讲到达时采取的行动。例如,当说话人还在说话时,自然的话轮转换行为就会发生,比如反向通道(嗯)和打断。在倾听时也会发生对可能参照物的眼动:个体递增地处理指令,在听到指令中的相关词语后立即对物体进行扫视眼动(Tanenhaus等人。1995);句子中出现的动词影响哪些对象是相关的工作我们以前已经表明,增量语法分析可以比非增量语法分析更快、更准确(Stoness等人。2005年。)此外,我们已经表明,在我们的试验床领域中,交互风格更强的语言的相对百分比也随着时间的推移而增加(Aist等人。2005年。)
Incremental Dialogue System Faster than and Preferred to its Nonincremental Counterpart Gregory Aist 1 (gregory.aist@asu.edu), James Allen 2 (james@cs.rochester.edu), Ellen Campana 2,3,4,5 (ecampana@bcs.rochester.edu), Carlos Gomez Gallo 2 (cgomez@cs.rochester.edu), Scott Stoness 2 (stoness@cs.rochester.edu), Mary Swift 2 (swift@cs.rochester.edu), and Michael K. Tanenhaus 3 (mtan@bcs.rochester.edu) Department of Computer Science and Engineering Arizona State University P.O. Box 878809 Tempe AZ 85287 Department of Computer Science University of Rochester P.O. Box 270226 Rochester NY 14627 Department of Brain and Cognitive Sciences University of Rochester P.O. Box 270268 Rochester, NY 14627 Abstract understanding; Arts, Media, and Engineering Program Arizona State University P.O. Box 878709 Tempe AZ 85287 Department of Psychology Arizona State University P.O. Box 871104 Tempe AZ 85287 brought into context, as determined by hearer eye fixations (Altmann and Kamide 1999). Other actions can also be taken based on partial utterances. Many different sources of knowledge are available for use in understanding. On the speech recognition side, commonly used sources of information include acoustics, phonetics and phonemics, lexical probability, and word order. In dialogue systems, additional sources of information often include syntax and semantics (both general and domain-specific.) There are also however some sources of information that are less frequently programmed. These include such linguistic information as morphology and prosody. Knowledge-based features are also available, such as world knowledge (triangles have three sides), domain knowledge (here there are two sizes of triangles), and task knowledge (the next step is to click on a small triangle.) There is also pragmatic information available from the visual context (there is a small triangle near the flag.) In this paper we discuss some of the progress we have made towards building methods for incremental understanding of spoken language by machines. We first discuss some of our and others’ related work in this area. We then discuss the testbed domain that we have been developing, and show some of the characteristics of human dialogue in the domain. We then discuss the incremental architecture that we have been developing, highlighting its differences from traditional architectures. Finally, we present an experimental evaluation of the performance of the system showing that incremental systems are both faster than and preferred to their nonincremental counterparts. Current dialogue systems generally operate in a pipelined, modular fashion on one complete utterance at a time. Evidence from human language understanding shows that human understanding operates incrementally and makes use of multiple sources of information during the parsing process, including traditionally “later” components such as pragmatics. In this paper we describe a spoken dialogue sys- tem that understands language incrementally, provides visual feedback on possible referents during the course of the user’s utterance, and allows for overlapping speech and actions. We further present findings from an empirical study showing that the resulting dialogue system is faster overall than its nonincremental counterpart. Furthermore, the incremental system is preferred to its nonincremental counterpart – beyond what is accounted for by factors such as speed and accuracy. These results indicate that successful incremental understanding systems will improve both performance and usability. Keywords: natural language systems; incremental processing. dialogue Introduction The standard model of natural language understanding for dialogue systems is pipelined, modular, and operates on complete utterances. By pipelined we mean that only one level of processing operates at a time, in a sequential manner. By modular, we mean that each level of processing depends only on the previous level. By complete utterances we mean that the system operates on one sentence at a time. There is, however, considerable evidence that human language processing is neither pipelined nor modular nor whole-utterance (Marslen-Wilson 1993). Evidence is converging from a variety of sources, including particularly actions taken while speech arrives. For example, natural turn-taking behavior such as backchanneling (uh-huh) and interruption occur while the speaker is still speaking. Eye movements to possible referents also occur while listening: individuals process instructions incrementally, making saccadic eye movements to objects right after hearing relevant words in the instruction (Tanenhaus et al. 1995); verbs appearing earlier in sentences affect which objects are Related Work We have previously shown that incremental parsing can be faster and more accurate than non-incremental parsing (Stoness et al. 2005.) In addition, we have shown that in our testbed domain the relative percentage of language that is of a more interactive style also increases over time (Aist et al. 2005.)