Interactive Evaluation of Dialog Track at DSTC9

Interactive Evaluation of Dialog Track at DSTC9
复制标题

DOI:
10.48550/arxiv.2207.14403
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Shikib Mehri;Yulan Feng;Carla Gordon;S. Alavi;D. Traum;M. Eskénazi
Shikib Mehri;Yulan Feng;Carla Gordon;S. Alavi;D. Traum;M. Eskénazi
中科院分区:
其他
文献类型:
--
作者:
Shikib Mehri;Yulan Feng;Carla Gordon;S. Alavi;D. Traum;M. Eskénazi

文献摘要

被引文献

相似文献

对话研究的最终目标是开发能够被真实用户在交互环境中有效使用的系统。为此,我们在第九届对话系统技术挑战赛上介绍了对话赛道的交互评估。该跟踪由两个子任务组成。第一个子任务涉及建立基于知识的反应生成模型。第二个子任务旨在通过在与真实用户交互的环境中对对话模型进行评估,将对话模型扩展到静态数据集之外。我们的跟踪要求参与者开发强大的响应生成模型,并探索将其扩展到与真实用户的来回互动的策略。从静态语料库到交互式评估的发展带来了独特的挑战,并促进了对开放领域对话系统的更彻底的评估。本文件概述了该轨道,包括方法和结果。此外,它还提供了有关如何最好地评估开放域对话模型的见解。
The ultimate goal of dialog research is to develop systems that can be effectively used in interactive settings by real users. To this end, we introduced the Interactive Evaluation of Dialog Track at the 9th Dialog System Technology Challenge. This track consisted of two sub-tasks. The first sub-task involved building knowledge-grounded response generation models. The second sub-task aimed to extend dialog models beyond static datasets by assessing them in an interactive setting with real users. Our track challenges participants to develop strong response generation models and explore strategies that extend them to back-and-forth interactions with real users. The progression from static corpora to interactive evaluation introduces unique challenges and facilitates a more thorough assessment of open-domain dialog systems. This paper provides an overview of the track, including the methodology and results. Furthermore, it provides insights into how to best evaluate open-domain dialog models.