JWSAN: Japanese word similarity and association norm

JWSAN: Japanese word similarity and association norm
复制标题

JWSAN:日语单词相似度和关联规范

DOI:
10.1007/s10579-021-09543-7
复制
发表时间:
2022
影响因子:
2.7
通讯作者:
Inohara Keisuke and Akira Utsumi
Inohara Keisuke and Akira Utsumi
中科院分区:
计算机科学4区
文献类型:
--
作者:
Inohara Keisuke and Akira Utsumi

文献摘要

相似文献

我们提出了一个新的日语数据集,日语单词相似性和关联规范(JWSAN),包括人类的相似性和关联的评分为2145个单词对,单词相似性和单词关联之间有明显的区别。人类语义记忆或心理词典的计算模型,如分布式语义模型,不仅要预测关联性,还要预测相似性。人们可以区分单词的相似性和关联性。然而,尽管SimLex-999数据集是公开的英语数据集,但没有日语相似性数据集在两种类型的单词相关性之间有明显的区别。JWSAN是第一个具有相似性和关联性评级的大型日语数据集,包含名词,动词和形容词词对。它的特点也是从足够数量的年龄和性别控制的评估员,通过基于网络的调查6450母语的日本人获得的相似性和关联评级的数据收集。此外,还研究了评分者的性别和年龄的影响,这些因素在过去很少考虑。该数据集可以作为改进日语分布式语义模型的基准。
We present a new Japanese dataset, Japanese Word Similarity and Association Norm (JWSAN), comprising human rating scores of similarity and association for 2145 word pairs, with a clear distinction between word similarity and word association. Computational models of human semantic memory or mental lexicon, such as distributed semantic models, must predict not only association but also similarity. People can distinguish between word similarity and association. However, although the SimLex-999 dataset is publicly available for English, there is no Japanese similarity dataset with a clear distinction between the two types of word relatedness. JWSAN is the first large Japanese dataset with similarity and association ratings, containing noun, verb, and adjective word pairs. It is also characterized by data collection from a sufficient number of age- and-gender-controlled assessors, with similarity and association ratings obtained via a web-based survey conducted of 6450 native speakers of Japanese. In addition, the effects of the gender and age of the raters were also examined; these factors were only given scant consideration in the past. This dataset can act as a benchmark for improving distributed semantic models in Japanese.