USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation

USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
复制标题

DOI:
10.18653/v1/2020.acl-main.64
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Shikib Mehri;M. Eskénazi
Shikib Mehri;M. Eskénazi
中科院分区:
其他
文献类型:
--
作者:
Shikib Mehri;M. Eskénazi

文献摘要

被引文献

相似文献

由于缺乏有意义的对话自动评价指标,阻碍了开放域对话的研究。标准的语言生成度量已被证明是无效的评估对话模型。为此,本文提出了USR,一个无监督和无参考的对话评价指标。USR是一种无参考指标,可训练无监督模型来衡量对话的几种理想质量。USR被证明与人类对Topical-Chat(回合级别:0.42,系统级别:1.0)和PersonaChat(回合级别:0.48和系统级别:1.0)的判断密切相关。USR还为对话框的几个理想属性提供了可解释的措施。
The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research. Standard language generation metrics have been shown to be ineffective for evaluating dialog models. To this end, this paper presents USR, an UnSupervised and Reference-free evaluation metric for dialog. USR is a reference-free metric that trains unsupervised models to measure several desirable qualities of dialog. USR is shown to strongly correlate with human judgment on both Topical-Chat (turn-level: 0.42, system-level: 1.0) and PersonaChat (turn-level: 0.48 and system-level: 1.0). USR additionally produces interpretable measures for several desirable properties of dialog.