A Comprehensive Analysis of PMI-based Models for Measuring Semantic Differences

A Comprehensive Analysis of PMI-based Models for Measuring Semantic Differences
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Taichi Aida;Mamoru Komachi;Toshinobu Ogiso;Hiroya Takamura;D. Mochihashi
Taichi Aida;Mamoru Komachi;Toshinobu Ogiso;Hiroya Takamura;D. Mochihashi
中科院分区:
其他
文献类型:
--
作者:
Taichi Aida;Mamoru Komachi;Toshinobu Ogiso;Hiroya Takamura;D. Mochihashi

文献摘要

相似文献

在语料库中检测具有语义差异的单词的任务主要由单词表示(如word 2 vec或BERT)来解决。然而,在语言学家和社会学家应用这些技术的真实的世界中,计算资源通常是有限的。在本文中,我们扩展了一个现有的同时优化的模型,可以在CPU上训练来执行这项任务。实验结果表明,扩展后的模型在英语语料库和SemEval-2020任务1以及日语中均取得了与强基线相当或上级的结果。此外,我们比较了每个模型的训练时间,并进行了全面的日语语料库分析。1
The task of detecting words with semantic differences across corpora is mainly addressed by word representations such as word2vec or BERT. However, in the real world where lin-guists and sociologists apply these techniques, computational resources are typically limited. In this paper, we extend an existing simultaneously optimized model that can be trained on CPU to perform this task. Experimental re-sults show that the extended models achieved comparable or superior results to strong base-lines in English corpora and SemEval-2020 Task 1, and also in Japanese. Furthermore, we compared the training time of each model and conducted a comprehensive analysis of Japanese corpora. 1