Exploring what is encoded in distributional word vectors: A neurobiologically motivated analysis

Exploring what is encoded in distributional word vectors: A neurobiologically motivated analysis
复制标题

探索分布式词向量中编码的内容:神经生物学动机分析

DOI:
10.1111/cogs.12844
复制
发表时间:
2020
期刊:
影响因子:
2.5
通讯作者:
Akira Utsumi
Akira Utsumi
中科院分区:
心理学3区
文献类型:
--
作者:
二井 昭佳;岡田 一天;Akira Utsumi

文献摘要

相似文献

分布式语义模型或词嵌入在认知建模和实际应用中的广泛使用是因为它们具有出色的表示词义的能力。然而,相对较少的努力已经探索什么类型的信息编码在分布式词向量。了解嵌入在词向量中的内部知识对于使用分布式语义模型进行认知建模非常重要。因此,在本文中,我们试图通过使用Binder等人进行的计算实验来识别编码在词向量中的知识。的(2016)基于神经生物学动机属性的特征概念表征。在实验中,使用神经网络和线性变换从基于文本的词向量预测这些概念向量,并在各种类型的信息中比较预测性能。分析表明,词向量对抽象信息的预测准确率普遍高于感知信息和时空信息,尤其是对认知信息和社会信息的预测准确率更高。情感信息也被成功地预测为抽象的话。这些结果表明,语言可以是一个主要的知识来源的抽象属性,他们支持最近的观点,强调语言的重要性,抽象概念。此外,我们表明,词向量可以捕获一些类型的感知和时空信息的具体概念和一些相关的词类别。这表明语言统计可以编码比通常预期的更多的感知知识。
The pervasive use of distributional semantic models or word embeddings for both cognitive modeling and practical application is because of their remarkable ability to represent the meanings of words. However, relatively little effort has been made to explore what types of information are encoded in distributional word vectors. Knowing the internal knowledge embedded in word vectors is important for cognitive modeling using distributional semantic models. Therefore, in this paper, we attempt to identify the knowledge encoded in word vectors by conducting a computational experiment using Binder et al.'s (2016) featural conceptual representations based on neurobiologically motivated attributes. In an experiment, these conceptual vectors are predicted from text‐based word vectors using a neural network and linear transformation, and prediction performance is compared among various types of information. The analysis demonstrates that abstract information is generally predicted more accurately by word vectors than perceptual and spatiotemporal information, and specifically, the prediction accuracy of cognitive and social information is higher. Emotional information is also found to be successfully predicted for abstract words. These results indicate that language can be a major source of knowledge about abstract attributes, and they support the recent view that emphasizes the importance of language for abstract concepts. Furthermore, we show that word vectors can capture some types of perceptual and spatiotemporal information about concrete concepts and some relevant word categories. This suggests that language statistics can encode more perceptual knowledge than often expected.