Investigating Evaluation of Open-Domain Dialogue Systems With Human Generated Multiple References

Investigating Evaluation of Open-Domain Dialogue Systems With Human Generated Multiple References
复制标题

DOI:
10.18653/v1/w19-5944
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Prakhar Gupta;Shikib Mehri;Tiancheng Zhao;Amy Pavel;M. Eskénazi;Jeffrey P. Bigham
Prakhar Gupta;Shikib Mehri;Tiancheng Zhao;Amy Pavel;M. Eskénazi;Jeffrey P. Bigham
中科院分区:
其他
文献类型:
--
作者:
Prakhar Gupta;Shikib Mehri;Tiancheng Zhao;Amy Pavel;M. Eskénazi;Jeffrey P. Bigham

文献摘要

被引文献

相似文献

本文的目的是通过多参考评估来缓解开放域对话系统自动评估的缺陷。现有的指标已被证明与人类判断的相关性较差,尤其是在开放域对话中。一种替代方法是收集人工标注用于评估,但这可能既昂贵又耗时。为了证明多参考评估的有效性,我们为DailyDialog的测试集增加了多个参考。一系列实验表明,使用多个参考提高了几个自动指标与人类对系统输出质量和多样性判断之间的相关性。
The aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multi-reference evaluation. Existing metrics have been shown to correlate poorly with human judgement, particularly in open-domain dialog. One alternative is to collect human annotations for evaluation, which can be expensive and time consuming. To demonstrate the effectiveness of multi-reference evaluation, we augment the test set of DailyDialog with multiple references. A series of experiments show that the use of multiple references results in improved correlation between several automatic metrics and human judgement for both the quality and the diversity of system output.