MultiWOZ 2.4: A Multi-Domain Task-Oriented Dialogue Dataset with Essential Annotation Corrections to Improve State Tracking Evaluation

MultiWOZ 2.4: A Multi-Domain Task-Oriented Dialogue Dataset with Essential Annotation Corrections to Improve State Tracking Evaluation
复制标题

DOI:
10.18653/v1/2022.sigdial-1.34
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Fanghua Ye;Jarana Manotumruksa;Emine Yilmaz
Fanghua Ye;Jarana Manotumruksa;Emine Yilmaz
中科院分区:
其他
文献类型:
--
作者:
Fanghua Ye;Jarana Manotumruksa;Emine Yilmaz

文献摘要

被引文献

相似文献

MultiWOZ 2.0数据集极大地促进了面向任务的对话系统的研究。然而,它的状态注释包含大量的噪声,这阻碍了对模型性能的正确评估。为了解决这个问题,人们投入了大量的精力来纠正注释。随后发布了三个改进版本(即MultiWOZ 2.1-2.3)。尽管如此,仍然存在大量不正确和不一致的注释。本文介绍了MultiWOZ 2.4,它对MultiWOZ 2.1的验证集和测试集中的注释进行了细化。训练集中的注释保持不变(与MultiWOZ 2.1相同),以获得鲁棒性和抗噪声的模型训练。我们在MultiWOZ 2.4上对8个最先进的对话状态跟踪模型进行了基准测试。它们都比MultiWOZ 2.1表现出更高的性能。
The MultiWOZ 2.0 dataset has greatly stimulated the research of task-oriented dialogue systems. However, its state annotations contain substantial noise, which hinders a proper evaluation of model performance. To address this issue, massive efforts were devoted to correcting the annotations. Three improved versions (i.e., MultiWOZ 2.1-2.3) have then been released. Nonetheless, there are still plenty of incorrect and inconsistent annotations. This work introduces MultiWOZ 2.4, which refines the annotations in the validation set and test set of MultiWOZ 2.1. The annotations in the training set remain unchanged (same as MultiWOZ 2.1) to elicit robust and noise-resilient model training. We benchmark eight state-of-the-art dialogue state tracking models on MultiWOZ 2.4. All of them demonstrate much higher performance than on MultiWOZ 2.1.