Evaluating Dialogue Response Generation Systems via Response Selection with Well-chosen False Candidates

Evaluating Dialogue Response Generation Systems via Response Selection with Well-chosen False Candidates
复制标题

DOI:
10.5715/jnlp.29.53
复制
发表时间:
2022
期刊:
Journal of Natural Language Processing
影响因子:
--
通讯作者:
Shiki Sato;Reina Akama;Hiroki Ouchi;Jun Suzuki;Kentaro Inui
Shiki Sato;Reina Akama;Hiroki Ouchi;Jun Suzuki;Kentaro Inui
中科院分区:
其他
文献类型:
--
作者:
Shiki Sato;Reina Akama;Hiroki Ouchi;Jun Suzuki;Kentaro Inui

文献摘要

相似文献

为开放域对话生成系统开发,可以以低成本验证每日系统改进的E(CID:11)。但是,通常用于自动响应生成评估的现有指标,例如双语评估研究(BLEU),与人类评估相结合。这种差的相关性源于对话的性质,即对输入上下文的几种可接受的反应。为了解决这个问题,我们专注于通过响应选择评估响应生成系统。在此任务中,对于给定的上下文,系统从一组响应候选者中选择适当的响应。由于系统只能选择特定(CID:12)C候选者,因此通过响应选择的评估可以减轻对话的上述性质的E(CID:11)。通常,虚假的响应候选者是从其他无关的对话中随机取样的,这导致了两个问题:(a)无关的假候选人和(b)被标记为假的可接受的话语。由于这些问题,一般响应选择测试集是不可靠的。因此,本文提出了一种使用精心挑选的假候选者构建响应选择测试集的方法。实验表明,与经常使用的自动评估指标(如BLEU)相比,通过精心挑选的假候选者进行响应选择评估系统与人类评估更加密切。
Developing for open-domain dialogue generation systems that can validate the e(cid:11)ects of daily system improvements at a low cost is necessary. However, existing metrics commonly used for automatic response generation evaluation, such as bilingual evaluation understudy (BLEU), cor-relate poorly with human evaluation. This poor correlation arises from the nature of dialogue, i.e., several acceptable responses to an input context. To address this issue, we focus on evaluating response generation systems via response selection. In this task, for a given context, systems select an appropriate response from a set of response candidates. Because the systems can only select speci(cid:12)c candidates, evaluation via response selection can mitigate the e(cid:11)ect of the above-mentioned nature of dialogue. Generally, false response candidates are randomly sampled from other unrelated dialogues, resulting in two issues: (a) unrelated false candidates and (b) acceptable utterances marked as false. General response selection test sets are unreliable owing to these issues. Thus, this paper proposes a method for constructing response selection test sets with well-chosen false candidates. Experiments demonstrate that evaluating systems via response selection with well-chosen false candidates correlates more strongly with human evaluation compared with commonly used automatic evaluation metrics such as BLEU.