A Comprehensive Assessment of Dialog Evaluation Metrics

A Comprehensive Assessment of Dialog Evaluation Metrics
复制标题

DOI:
10.18653/v1/2021.eancs-1.3
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yi-Ting Yeh;M. Eskénazi;Shikib Mehri
Yi-Ting Yeh;M. Eskénazi;Shikib Mehri
中科院分区:
其他
文献类型:
--
作者:
Yi-Ting Yeh;M. Eskénazi;Shikib Mehri

文献摘要

相似文献

自动评价指标是对话系统研究的重要组成部分。已知标准语言评估度量对于评估对话是无效的。因此,最近的研究提出了一些新的,对话特定的指标,更好地与人类的判断。由于研究的快速发展,许多指标已经在不同的数据集上进行了评估,目前还没有时间对它们进行系统的比较。为此,本文提供了一个全面的评估最近提出的对话评价指标的一些数据集。在本文中,23个不同的自动评估指标进行了评估,对10个不同的数据集。此外,在不同的环境中评估这些指标,以更好地确定其各自的长处和短处。这一全面的评估提供了几个关于对话评估指标的要点。它还就如何最好地评估评价指标提出了建议,并为今后的工作指明了有希望的方向。
Automatic evaluation metrics are a crucial component of dialog systems research. Standard language evaluation metrics are known to be ineffective for evaluating dialog. As such, recent research has proposed a number of novel, dialog-specific metrics that correlate better with human judgements. Due to the fast pace of research, many of these metrics have been assessed on different datasets and there has as yet been no time for a systematic comparison between them. To this end, this paper provides a comprehensive assessment of recently proposed dialog evaluation metrics on a number of datasets. In this paper, 23 different automatic evaluation metrics are evaluated on 10 different datasets. Furthermore, the metrics are assessed in different settings, to better qualify their respective strengths and weaknesses. This comprehensive assessment offers several takeaways pertaining to dialog evaluation metrics in general. It also suggests how to best assess evaluation metrics and indicates promising directions for future work.