Automatic Word Sense Discrimination

Automatic Word Sense Discrimination
复制标题

DOI:
--
复制
发表时间:
1998-03
期刊:
Comput. Linguistics
影响因子:
--
通讯作者:
Hinrich Schütze
Hinrich Schütze
中科院分区:
其他
文献类型:
--
作者:
Hinrich Schütze

文献摘要

被引文献

相似文献

本文提出了一种基于聚类的消歧算法--上下文组区分算法。意义被解释为歧义单词的相似上下文的组(或群)。词、上下文和意义在词空间中表示,词空间是一个高维的实值空间,在这个空间中,接近度对应于语义相似性。词空间中的相似性基于二阶共现:如果歧义单词的两个标记(或上下文)与它们共同出现的单词在训练语料库中与相似单词一起出现,则将它们分配到相同的义簇。该算法在训练和应用中都是自动的和无监督的:意义是从没有标记的训练实例或其他外部知识源的语料库中归纳出来的。通过对自然歧义词和人工歧义词的实验,证明了该方法具有较好的上下文区分性能。
This paper presents context-group discrimination, a disambiguation algorithm based on clustering. Senses are interpreted as groups (or clusters) of similar contexts of the ambiguous word. Words, contexts, and senses are represented in Word Space, a high-dimensional, real-valued space in which closeness corresponds to semantic similarity. Similarity in Word Space is based on second-order co-occurrence: two tokens (or contexts) of the ambiguous word are assigned to the same sense cluster if the words they co-occur with in turn occur with similar words in a training corpus. The algorithm is automatic and unsupervised in both training and application: senses are induced from a corpus without labeled training instances or other external knowledge sources. The paper demonstrates good performance of context-group discrimination for a sample of natural and artificial ambiguous words.