A review of evaluation techniques for social dialogue systems

A review of evaluation techniques for social dialogue systems
复制标题

DOI:
10.1145/3139491.3139504
复制
发表时间:
2017-08
期刊:
Proceedings of the 1st ACM SIGCHI International Workshop on Investigating Social Interactions with Artificial Agents
影响因子:
--
通讯作者:
A. C. Curry;H. Hastie;Verena Rieser
A. C. Curry;H. Hastie;Verena Rieser
中科院分区:
其他
文献类型:
--
作者:
A. C. Curry;H. Hastie;Verena Rieser

文献摘要

被引文献

相似文献

与目标导向的对话相比,社会对话没有明确的任务成功衡量标准。因此,对这些系统的评估非常困难。在本文中,我们回顾了当前的评估方法,重点关注自动指标。我们得出的结论是,基于回合的指标通常会忽略上下文,并且不会考虑多个回复有效的事实,而对话结束奖励主要是手工制作的。两者都缺乏人类认知的基础。
In contrast with goal-oriented dialogue, social dialogue has no clear measure of task success. Consequently, evaluation of these systems is notoriously hard. In this paper, we review current evaluation methods, focusing on automatic metrics. We conclude that turn-based metrics often ignore the context and do not account for the fact that several replies are valid, while end-of-dialogue rewards are mainly hand-crafted. Both lack grounding in human perceptions.