Latent Semantic Word Sense Disambiguation Using Global Co-occurrence Information

Latent Semantic Word Sense Disambiguation Using Global Co-occurrence Information
复制标题

DOI:
10.5121/csit.2014.4240
复制
发表时间:
2014-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Minoru Sasaki
Minoru Sasaki
中科院分区:
其他
文献类型:
--
作者:
Minoru Sasaki

文献摘要

相似文献

本文提出了一种新的基于NMF的基于全局共现信息的词义消歧方法。当我计算依赖关系矩阵时,现有的方法往往从一个小的训练集产生非常稀疏的共生矩阵。因此,NMF算法有时不能收敛到期望的解。为了获得大量的共现关系,本文提出了利用整个训练集中词语特征之间依存关系的共现频率。这使得我们能够解决数据稀疏问题,并提取出更有效的潜在特征。为了评估词义消歧方法的效率,我进行了一些实验,并与两种基线方法的结果进行了比较。实验结果表明,与所有的基线方法相比,该方法对词义消歧是有效的。此外,该方法通过分析全局共现信息,有效地获得了稳定的效果。
In this paper, I propose a novel word sense disambiguation method based on the global co-occurrence information using NMF. When I calculate the dependency relation matrix, the existing method tends to produce very sparse co-occurrence matrix from a small training set. Therefore, the NMF algorithm sometimes does not converge to desired solutions. To obtain a large number of co-occurrence relations, I propose to use co-occurrence frequencies of dependency relations between word features in the whole training set. This enables us to solve data sparseness problem and induce more effective latent features. To evaluate the efficiency of the method of word sense disambiguation, I make some experiments to compare with the result of the two baseline methods. The results of the experiments show this method is effective for word sense disambiguation in comparison with the all baseline methods. Moreover, the proposed method is effective for obtaining a stable effect by analyzing the global co-occurrence information.