Evaluating Dialogue Response Generation Systems via Response Selection with Well-chosen False Candidates
Evaluating Dialogue Response Generation Systems via Response Selection with Well-chosen False Candidates
复制标题
DOI:
10.5715/jnlp.29.53
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Shiki Sato;Reina Akama;Hiroki Ouchi;Jun Suzuki;Kentaro Inui
中科院分区:
文献类型:
--
作者:
Shiki Sato;Reina Akama;Hiroki Ouchi;Jun Suzuki;Kentaro Inui
Developing for open-domain dialogue generation systems that can validate the e(cid:11)ects of daily system improvements at a low cost is necessary. However, existing metrics commonly used for automatic response generation evaluation, such as bilingual evaluation understudy (BLEU), cor-relate poorly with human evaluation. This poor correlation arises from the nature of dialogue, i.e., several acceptable responses to an input context. To address this issue, we focus on evaluating response generation systems via response selection. In this task, for a given context, systems select an appropriate response from a set of response candidates. Because the systems can only select speci(cid:12)c candidates, evaluation via response selection can mitigate the e(cid:11)ect of the above-mentioned nature of dialogue. Generally, false response candidates are randomly sampled from other unrelated dialogues, resulting in two issues: (a) unrelated false candidates and (b) acceptable utterances marked as false. General response selection test sets are unreliable owing to these issues. Thus, this paper proposes a method for constructing response selection test sets with well-chosen false candidates. Experiments demonstrate that evaluating systems via response selection with well-chosen false candidates correlates more strongly with human evaluation compared with commonly used automatic evaluation metrics such as BLEU.