MultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselines

MultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselines
复制标题

DOI:
10.18653/v1/2020.nlp4convai-1.13
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Xiaoxue Zang;Abhinav Rastogi;Jianguo Zhang;Jindong Chen
Xiaoxue Zang;Abhinav Rastogi;Jianguo Zhang;Jindong Chen
中科院分区:
其他
文献类型:
--
作者:
Xiaoxue Zang;Abhinav Rastogi;Jianguo Zhang;Jindong Chen

文献摘要

被引文献

相似文献

MultiWOZ是一个著名的面向任务的对话数据集,包含跨越8个域的10,000多个带注释的对话。它被广泛用作对话状态跟踪的基准。然而,最近的著作报道了对话状态注释中存在大量噪音。MultiWOZ 2.1识别并修复了许多此类错误注释和用户发言,从而改进了此数据集的版本。这项工作引入了MultiWOZ 2.2,它是该数据集的另一个改进版本。首先,我们在MultiWOZ 2.1上识别并修复了17.3%的对话状态标注错误。其次,我们通过禁止具有大量可能值(例如餐厅名称、预订时间)的槽的词汇来重新定义本体。此外,我们为这些槽引入了槽跨度注释,以在最近的模型中对它们进行标准化,这些模型以前使用定制的字符串匹配启发式算法来生成它们。我们还在校正后的数据集上对几个最先进的对话状态跟踪模型进行了基准测试,以便于未来工作的比较。最后,我们讨论了有助于避免注释错误的对话数据收集的最佳实践。
MultiWOZ is a well-known task-oriented dialogue dataset containing over 10,000 annotated dialogues spanning 8 domains. It is extensively used as a benchmark for dialogue state tracking. However, recent works have reported presence of substantial noise in the dialogue state annotations. MultiWOZ 2.1 identified and fixed many of these erroneous annotations and user utterances, resulting in an improved version of this dataset. This work introduces MultiWOZ 2.2, which is a yet another improved version of this dataset. Firstly, we identify and fix dialogue state annotation errors across 17.3% of the utterances on top of MultiWOZ 2.1. Secondly, we redefine the ontology by disallowing vocabularies of slots with a large number of possible values (e.g., restaurant name, time of booking). In addition, we introduce slot span annotations for these slots to standardize them across recent models, which previously used custom string matching heuristics to generate them. We also benchmark a few state of the art dialogue state tracking models on the corrected dataset to facilitate comparison for future work. In the end, we discuss best practices for dialogue data collection that can help avoid annotation errors.