Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness

Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
复制标题

DOI:
10.18653/v1/2020.emnlp-main.68
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Reina Akama;Sho Yokoi;Jun Suzuki;Kentaro Inui
Reina Akama;Sho Yokoi;Jun Suzuki;Kentaro Inui
中科院分区:
其他
文献类型:
--
作者:
Reina Akama;Sho Yokoi;Jun Suzuki;Kentaro Inui

文献摘要

相似文献

大规模的对话数据集最近已经可以用于训练神经对话代理。然而,这些数据集已被报告包含不可忽略数量的不可接受的话语对。在本文中,我们提出了一种方法来评分的话语对的连接性和相关性的质量。建议的评分方法是基于对话和语言学研究社区广泛共享的发现而设计的。我们证明,它有一个相对较好的相关性与人类的判断对话质量。此外,该方法被应用于过滤出潜在的不可接受的话语对从大规模的噪声对话语料库,以确保其质量。我们通过实验证实,通过所提出的方法过滤的训练数据提高了响应生成中神经对话代理的质量。
Large-scale dialogue datasets have recently become available for training neural dialogue agents. However, these datasets have been reported to contain a non-negligible number of unacceptable utterance pairs. In this paper, we propose a method for scoring the quality of utterance pairs in terms of their connectivity and relatedness. The proposed scoring method is designed based on findings widely shared in the dialogue and linguistics research communities. We demonstrate that it has a relatively good correlation with the human judgment of dialogue quality. Furthermore, the method is applied to filter out potentially unacceptable utterance pairs from a large-scale noisy dialogue corpus to ensure its quality. We experimentally confirm that training data filtered by the proposed method improves the quality of neural dialogue agents in response generation.