D-Score: Holistic Dialogue Evaluation Without Reference

D-Score: Holistic Dialogue Evaluation Without Reference
复制标题

D-Score:无参考的整体对话评估

DOI:
10.1109/taslp.2021.3074012
复制
发表时间:
2021
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Haizhou Li
Haizhou Li
中科院分区:
--
文献类型:
--
作者:
Chen Zhang;Grandee Lee;L. F. D’Haro;Haizhou Li

文献摘要

被引文献

相似文献

在艺术体操比赛中,难度分或D分用于评定成绩。从零开始,运动员从不同方面获得积分,如构图要求,难度和动作之间的联系。最终得分是由各项性能指标的质量组成。同样,在评估对话响应时,人类法官通常遵循一些标准,其中语言流畅性,上下文连贯性,逻辑一致性和语义适当性是首要的。在本文中,我们提出了一个自动对话评估框架称为D-评分,类似于体操的方式进行评估。根据上述四个人类判断标准,我们设计了一系列评估任务,并在多任务学习框架下对其进行建模。所提出的框架,不依赖于任何人写的参考,学会欣赏人与人之间的对话的整体质量,通过一个表示,是由所有的任务共享,而不会过度拟合到个别的任务域。我们通过对三个对话评估数据集(其中两个来自过去的DSTC系列)进行与人类判断的综合相关性分析来评估D分数,并以最先进的基线为基准。D评分不仅在系统级斯皮尔曼相关性方面大幅优于最佳基线,而且也是迈向可解释对话评分的重要一步。
In artistic gymnastics, difficulty score or D-score is used for judging performance. Starting from zero, an athlete earns points from different aspects such as composition requirement, difficulty, and connection between moves. The final score is a composition of the quality of various performance indicators. Similarly, when evaluating dialogue responses, human judges generally follow a number of criteria, among which language fluency, context coherence, logical consistency, and semantic appropriateness are on top of the agenda. In this paper, we propose an automatic dialogue evaluation framework called D-score that resembles the way gymnastics is evaluated. Following the four human judging criteria above, we devise a range of evaluation tasks and model them under a multi-task learning framework. The proposed framework, without relying on any human-written reference, learns to appreciate the overall quality of human-human conversations through a representation that is shared by all tasks without over-fitting to individual task domain. We evaluate D-score by performing comprehensive correlation analyses with human judgement on three dialogue evaluation datasets, among which two are from past DSTC series, and benchmark against state-of-the-art baselines. D-score not only outperforms the best baseline by a large margin in terms of system-level Spearman correlation but also represents an important step towards explainable dialogue scoring.