GCDF1: A Goal- and Context- Driven F-Score for Evaluating User Models

GCDF1: A Goal- and Context- Driven F-Score for Evaluating User Models
复制标题

DOI:
10.18653/v1/2021.eancs-1.2
复制
发表时间:
2021
期刊:
The First Workshop on Evaluations and Assessments of Neural Conversation Systems
影响因子:
--
通讯作者:
Alexandru Coca;Bo-Hsiang Tseng;B. Byrne
Alexandru Coca;Bo-Hsiang Tseng;B. Byrne
中科院分区:
其他
文献类型:
--
作者:
Alexandru Coca;Bo-Hsiang Tseng;B. Byrne

文献摘要

相似文献

对话系统与模拟用户交互的评估已经被提出来改进回合级别的、基于语料库的度量,这些度量只能评估在语料库中遇到的测试用例,而不能测量系统维持多回合交互的能力。最近,很少强调自动评估用户模型本身的质量,所以除非与人类研究的相关性进行测量,基于用户模型的评估的可靠性是未知的。我们提出了GCDF 1,一个简单而有效的措施之间的语义级对话的质量目标驱动的用户代理和系统代理。与以前的方法相比,我们在对话级别测量F分数,并考虑用户和系统行为,以提高召回率和精度估计。我们通过提供丰富的层次结构与测试数据和工具中存在的对话模式的信息,以有效地查询生成的对话,从而促进分数的解释。我们应用我们的框架来评估Convlab2用户模型的性能和弱点。
The evaluation of dialogue systems in interaction with simulated users has been proposed to improve turn-level, corpus-based metrics which can only evaluate test cases encountered in a corpus and cannot measure system’s ability to sustain multi-turn interactions. Recently, little emphasis was put on automatically assessing the quality of the user model itself, so unless correlations with human studies are measured, the reliability of user model based evaluation is unknown. We propose GCDF1, a simple but effective measure of the quality of semantic-level conversations between a goal-driven user agent and a system agent. In contrast with previous approaches we measure the F-score at dialogue level and consider user and system behaviours to improve recall and precision estimation. We facilitate scores interpretation by providing a rich hierarchical structure with information about conversational patterns present in the test data and tools to efficiently query the conversations generated. We apply our framework to assess the performance and weaknesses of a Convlab2 user model.