Placement of Nouns in a Multi-Dimensional Space Based on Words' Cooccurrency

Placement of Nouns in a Multi-Dimensional Space Based on Words' Cooccurrency
复制标题

基于词共现的多维空间名词放置

DOI:
10.1527/tjsai.19.1
复制
发表时间:
2004
影响因子:
--
通讯作者:
T. Hitaka
T. Hitaka
中科院分区:
--
文献类型:
--
作者:
Yoichi Tomiura;Shosaku Tanaka;T. Hitaka

文献摘要

被引文献

相似文献

词之间的语义相似度(或距离)是自然语言处理的基本知识之一。在多维空间中,已有一些基于词向量的相似性(或距离)测量的研究。在这些研究中,从语料库中的词的共现或词典中的词的参考关系中得到词的高维特征向量,然后通过主成分分析等方法从特征向量中计算出词向量。本文提出了一种基于语料库中词共现的多维空间名词放置方法。该方法不使用单词的高维特征向量,而是基于“在关系f中与单词w同时出现的名词对应的向量在多维空间中构成一组”的思想。虽然该方法得到的词向量不能反映名词的全部含义,但用词向量定义的名词之间的语义相似度(或距离)适合于基于实例的消歧方法。
The semantic similarity (or distance) between words is one of the basic knowledge in Natural Language Processing. There have been several previous studies on measuring the similarity (or distance) based on word vectors in a multi-dimensional space. In those studies, high dimensional feature vectors of words are made from words' cooccurrence in a corpus or from reference relation in a dictionary, and then the word vectors are calculated from the feature vectors through the method like principal component analysis. This paper proposes a new placement method of nouns into a multi-dimensional space based on words' cooccurrence in a corpus. The proposed method doesn't use the high dimensional feature vectors of words, but is based on the idea that ``vectors corresponding to nouns which cooccur with a word w in a relation f constitute a group in the multi-dimensional space''. Although the whole meaning of nouns isn't reflected in the word vectors obtained by the pro posed method, the semantic similarity (or distance) between nouns defined with the word vectors is proper for an example-based disambiguation method.