Creating a Japanese Dialogue Corpus with Multi-level Topic Analysis

Creating a Japanese Dialogue Corpus with Multi-level Topic Analysis
复制标题

DOI:
10.1109/icnlp55136.2022.00065
复制
发表时间:
2022-03
期刊:
2022 4th International Conference on Natural Language Processing (ICNLP)
影响因子:
--
通讯作者:
Yuma Komoto;Xin Kang;F. Ren
Yuma Komoto;Xin Kang;F. Ren
中科院分区:
其他
文献类型:
--
作者:
Yuma Komoto;Xin Kang;F. Ren

文献摘要

相似文献

随着自然语言理解和生成技术的深入研究,生成式对话系统已成为近年来的研究热点。然而,大多数的研究都集中在那些广泛使用的语言,如英语和汉语,收集了大量的对话语料库,并对其进行了深入的分析,以训练对话生成模型。此外,在这些语料库中保持人类语言的多样性和自发性对于构建对话系统来说是一个挑战。本文提出了一种利用Twitter上发布的对话构建大型日语对话语料库的方法,并通过分析语料库中的词簇与对话簇在同一语义空间中的相似度,自动标注对话级和话语级话题标签及其相应的概率分数。这种方法的优点是,它不需要人类工作人员花费昂贵的时间和精力来模仿对话和注释标签,这对于那些不太广泛使用的语言(如日语)特别有用。我们比较了四种过滤设置对语料库创建和主题注释的话语长度下界的影响,并报告了基于人工评估的话语长度对对话语料库质量的影响。基于该语料库,我们进一步提出了两个基于主题的对话生成任务,即下一个响应主题预测任务和基于下一个主题的响应生成任务。日语对话语料库可在GitHub1上获得。
The study of generative dialogue systems has become a hotspot with the recently well-studied natural language understanding and generation techniques. However, most works have been focusing on those widely used languages, such as English and Chinese, with huge dialogue corpora being collected and thoroughly analyzed for training the dialogue generation models. In addition, retaining the human-like diversity and spontaneity of speech in these corpora is challenging for building dialogue systems. In this paper, we propose a method to build a lage Japanese dialogue corpus by using the conversations posted on Twitter and to annotate the dialogue- and utterance-level topic labels and the corresponding probabilistic scores automatically by analyzing the similarity between the word clusters and the dialogue clusters of the corpus in the same semantic space. The advantage of this method is that it does not require the expensive time and effort of human workers for mimicking dialogues and annotating labels, which is specifically useful for those less widely used languages, such as Japanese. We compare four filtering settings with respect to the lower-bound of utterance length for corpus creation and topic annotation and report the effect of the utterance length to the quality of our dialogue corpus, based on a manual evaluation. Based on this corpus, we further propose two topic-based dialogue generation tasks, that is, the next-response-topic prediction task and the next-topic-based response generation task. The Japanese dialogue corpus is available on GitHub1.