A Model for Word Sense Disambiguation

A Model for Word Sense Disambiguation
复制标题

词义消歧模型

DOI:
10.30019/ijclclp.199908.0001
复制
发表时间:
1999
期刊:
International Journal of Computational Linguistics and Chinese Language Processing
影响因子:
--
通讯作者:
C. Huang
C. Huang
中科院分区:
--
文献类型:
--
作者:
Juan;C. Huang

文献摘要

被引文献

相似文献

词义消歧是自然语言处理中最困难的问题之一。本文提出了一种将叙词表的结构语义空间映射到多维实值向量空间的模型,并给出了基于该映射的词义消歧方法。该模型采用无监督学习的方法获取消歧知识,不仅节省了大量的人工劳动,而且实现了对大量实词的词义标注。首先,本文利用汉语叙词表《词林》和大规模语料库构建语义空间结构。在此基础上,提出了一个动态排歧模型,该模型根据歧义词在各个可能类别中的语义向量对歧义词进行排歧。为了解决数据稀疏性问题,提出了一种增强模型鲁棒性的方法。测试结果表明,该模型具有较好的性能,也可用于其他语言。
Word sense disambiguation is one of the most difficult problems in natural language processing. This paper puts forward a model for mapping a structural semantic space from a thesaurus into a multi-dimensional, real-valued vector space and gives a word sense disambiguation method based on this mapping. The model, which uses an unsupervised learning method to acquire the disambiguation knowledge, not only saves extensive manual work, but also realizes the sense tagging of a large number of content words. Firstly, a Chinese thesaurus Cilin and a very large-scale corpus are used to construct the structure of the semantic space. Then, a dynamic disambiguation model is developed to disambiguate an ambiguous word according to the vectors of monosemous words in each of its possible categories. In order to resolve the problem of data sparseness, a method is proposed to make the model more robust. Testing results show that the model has relatively good performance and can also be used for other languages.