A tractable hybrid DDN–POMDP approach to affective dialogue modeling for probabilistic frame-based dialogue systems

A tractable hybrid DDN–POMDP approach to affective dialogue modeling for probabilistic frame-based dialogue systems
复制标题

DOI:
10.1017/s1351324908005032
复制
发表时间:
2008-10
影响因子:
2.5
通讯作者:
Trung H. Bui;M. Poel;A. Nijholt;J. Zwiers
Trung H. Bui;M. Poel;A. Nijholt;J. Zwiers
中科院分区:
计算机科学3区
文献类型:
--
作者:
Trung H. Bui;M. Poel;A. Nijholt;J. Zwiers

文献摘要

被引文献

相似文献

摘要我们提出了一种新的方法来开发一个易于处理的情感对话模型的概率框架为基础的对话系统。基于部分可观察马尔可夫决策过程(POMDP)和动态决策网络(DDN)技术的情感对话模型由两个主要部分组成:槽级对话管理器和全局对话管理器。它有两个新特点:(1)能够处理大量时隙,以及(2)能够在导出自适应对话策略时考虑用户情感状态的某些方面。我们实现的原型对话管理器可以处理数百个插槽,其中每个插槽可能有数百个值。我们的方法是通过在危机管理领域的路线导航的例子说明。我们进行了各种实验来评估我们的方法,并将其与近似POMDP技术和手工制作的政策进行比较。实验结果表明,DDN-POMDP策略优于三个手工策略时,用户的动作错误是由压力以及观察误差增加。此外,一步前瞻DDN-POMDP政策优化其内部奖励后的性能接近最先进的近似POMDP同行。
Abstract We propose a novel approach to developing a tractable affective dialogue model for probabilistic frame-based dialogue systems. The affective dialogue model, based on Partially Observable Markov Decision Process (POMDP) and Dynamic Decision Network (DDN) techniques, is composed of two main parts: the slot-level dialogue manager and the global dialogue manager. It has two new features: (1) being able to deal with a large number of slots and (2) being able to take into account some aspects of the user's affective state in deriving the adaptive dialogue strategies. Our implemented prototype dialogue manager can handle hundreds of slots, where each individual slot might have hundreds of values. Our approach is illustrated through a route navigation example in the crisis management domain. We conducted various experiments to evaluate our approach and to compare it with approximate POMDP techniques and handcrafted policies. The experimental results showed that the DDN–POMDP policy outperforms three handcrafted policies when the user's action error is induced by stress as well as when the observation error increases. Further, performance of the one-step look-ahead DDN–POMDP policy after optimizing its internal reward is close to state-of-the-art approximate POMDP counterparts.