TyDiP: A Dataset for Politeness Classification in Nine Typologically Diverse Languages

TyDiP: A Dataset for Politeness Classification in Nine Typologically Diverse Languages
复制标题

TyDiP:九种不同语言的礼貌分类数据集

DOI:
10.48550/arxiv.2211.16496
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Eunsol Choi
Eunsol Choi
中科院分区:
--
文献类型:
--
作者:
A. Srinivasan;Eunsol Choi

文献摘要

被引文献

相似文献

我们研究礼貌现象在九个类型不同的语言。礼貌是交际的一个重要方面,有时被认为是特定的文化,但现有的计算语言学研究仅限于英语。我们创建了TyDiP,这是一个包含三种礼貌注释的数据集,每种语言有500个例子,总共有4.5K个例子。我们评估了多语言模型可以识别礼貌水平的程度-它们显示出相当强大的零镜头传输能力,但显着低于估计的人类准确性。我们进一步研究了通过自动翻译和词汇归纳将英语礼貌策略词典映射到九种语言中,分析每种策略的影响是否在不同语言中保持一致。最后,我们通过迁移实验对正式与礼貌之间的复杂关系进行了实证研究。我们希望我们的数据集能够支持各种研究问题和应用,从评估多语言模型到构建礼貌的多语言代理。
We study politeness phenomena in nine typologically diverse languages. Politeness is an important facet of communication and is sometimes argued to be cultural-specific, yet existing computational linguistic study is limited to English. We create TyDiP, a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples. We evaluate how well multilingual models can identify politeness levels -- they show a fairly robust zero-shot transfer ability, yet fall short of estimated human accuracy significantly. We further study mapping the English politeness strategy lexicon into nine languages via automatic translation and lexicon induction, analyzing whether each strategy's impact stays consistent across languages. Lastly, we empirically study the complicated relationship between formality and politeness through transfer experiments. We hope our dataset will support various research questions and applications, from evaluating multilingual models to constructing polite multilingual agents.