SuperSim: a test set for word similarity and relatedness in Swedish

SuperSim: a test set for word similarity and relatedness in Swedish
复制标题

SuperSim:瑞典语单词相似性和相关性的测试集

DOI:
--
复制
发表时间:
2021
期刊:
Nordic Conference of Computational Linguistics
影响因子:
--
通讯作者:
Nina Tahmasebi
Nina Tahmasebi
中科院分区:
--
文献类型:
--
作者:
Simon Hengchen;Nina Tahmasebi

文献摘要

参考文献

被引文献

相似文献

众所周知,语言模型很难评估。我们发布了SuperSim,这是一个针对瑞典人的大规模相似性和关联性测试集,带有专家的人类判断。测试集由1360个词对组成,由五个注释员独立判断关联度和相似度。我们评估了在两个独立的瑞典数据集(即瑞典Gigaword语料库和瑞典维基百科转储)上训练的三种不同模型(word2vec、fast Text和Glove),以提供未来比较的基线。我们将发布带有完整注释的测试集、代码、模型和数据。
Language models are notoriously difficult to evaluate. We release SuperSim, a large-scale similarity and relatedness test set for Swedish built with expert human judgements. The test set is composed of 1,360 word-pairs independently judged for both relatedness and similarity by five annotators. We evaluate three different models (Word2Vec, fastText, and GloVe) trained on two separate Swedish datasets, namely the Swedish Gigaword corpus and a Swedish Wikipedia dump, to provide a baseline for future comparison. We will release the fully annotated test set, code, models, and data.
历时用法相关性(DURel):词汇语义变化的注释框架
DOI: 10.18653/v1/n18-2027
发表时间: 2018
期刊:
影响因子: --
作者:
Schlechtweg;Dominik;Sabine Schulte im Walde ;Stefanie Eckmann
通讯作者: Stefanie Eckmann