Automatic Learning and Evaluation of User-Centered Objective Functions for Dialogue System Optimisation

Automatic Learning and Evaluation of User-Centered Objective Functions for Dialogue System Optimisation
复制标题

用于对话系统优化的以用户为中心的目标函数的自动学习和评估

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Oliver Lemon
Oliver Lemon
中科院分区:
--
文献类型:
--
作者:
Verena Rieser;Oliver Lemon

文献摘要

参考文献

被引文献

相似文献

建立对话系统的最终目的是满足真实用户的需求,但对话策略的质量保证是一个不容忽视的问题。应用的评估指标和由此产生的设计原则通常是模糊的,通过反复试验而出现,并且高度依赖于上下文。本文介绍了为系统设计获取可靠目标函数的数据驱动方法。特别是,我们测试从绿野仙踪(WOZ)数据中获得的目标函数是否是对真实用户偏好的有效估计。我们在WOZ研究获得的模型和与真实用户测试时获得的模型之间的测试-重新测试比较中对此进行了测试。我们可以表明,尽管与初始数据的拟合度较低,但从WOZ数据获得的目标函数为自动对话评估做出了准确的预测,并且,当使用这些预测自动优化策略时,从误差分析中可以清楚地看出,相对于简单地模仿数据的策略的改进。
The ultimate goal when building dialogue systems is to satisfy the needs of real users, but quality assurance for dialogue strategies is a non-trivial problem. The applied evaluation metrics and resulting design principles are often obscure, emerge by trial-and-error, and are highly context dependent. This paper introduces data-driven methods for obtaining reliable objective functions for system design. In particular, we test whether an objective function obtained from Wizard-of-Oz (WOZ) data is a valid estimate of real users preferences. We test this in a test-retest comparison between the model obtained from the WOZ study and the models obtained when testing with real users. We can show that, despite a low fit to the initial data, the objective function obtained from WOZ data makes accurate predictions for automatic dialogue evaluation, and, when automatically optimising a policy using these predictions, the improvement over a strategy simply mimicking the data becomes clear from an error analysis.
DOI: 10.1162/coli.2008.07-028-r2-05-82
发表时间: 2008-12
影响因子: 9.3
作者:
James Henderson;Oliver Lemon;Kallirroi Georgila
通讯作者: James Henderson;Oliver Lemon;Kallirroi Georgila