Investigating Word-Class Distributions in Word Vector Spaces

Investigating Word-Class Distributions in Word Vector Spaces
复制标题

DOI:
10.18653/v1/2020.acl-main.337
复制
发表时间:
2020-07
期刊:
影响因子:
1.4
通讯作者:
Ryohei Sasano;A. Korhonen
Ryohei Sasano;A. Korhonen
中科院分区:
--
文献类型:
--
作者:
Ryohei Sasano;A. Korhonen

文献摘要

相似文献

本文提出了一种研究属于某个词类的词向量在预先训练的词向量空间中的分布。为此,我们对分布做了几个假设,对分布进行了相应的建模,并通过比较每个模型的优度来验证每个假设。具体来说,我们考虑了两种类型的词类-动词的直接对象的语义类和同义词库中的语义类-并试图建立模型,正确估计向量空间中的词是给定词类的成员的可能性。我们在选择偏好和WordNet数据集上的结果表明,基于质心的模型将无法实现足够好的性能,分布的几何形状和子组的存在将产生有限的影响,并且需要考虑负面情况才能对分布进行充分的建模。我们进一步研究了每个模型计算的分数与隶属度之间的关系,发现基于判别学习的模型在寻找类的边界方面表现最好,而基于正负实例之间偏移的模型在确定隶属度方面表现最好。
This paper presents an investigation on the distribution of word vectors belonging to a certain word class in a pre-trained word vector space. To this end, we made several assumptions about the distribution, modeled the distribution accordingly, and validated each assumption by comparing the goodness of each model. Specifically, we considered two types of word classes – the semantic class of direct objects of a verb and the semantic class in a thesaurus – and tried to build models that properly estimate how likely it is that a word in the vector space is a member of a given word class. Our results on selectional preference and WordNet datasets show that the centroid-based model will fail to achieve good enough performance, the geometry of the distribution and the existence of subgroups will have limited impact, and also the negative instances need to be considered for adequate modeling of the distribution. We further investigated the relationship between the scores calculated by each model and the degree of membership and found that discriminative learning-based models are best in finding the boundaries of a class, while models based on the offset between positive and negative instances perform best in determining the degree of membership.