CFILT: Resource Conscious Approaches for All-Words Domain Specific WSD

CFILT: Resource Conscious Approaches for All-Words Domain Specific WSD
复制标题

CFILT:针对全字域特定 WSD 的资源意识方法

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
P. Bhattacharyya
P. Bhattacharyya
中科院分区:
--
文献类型:
--
作者:
A. Kulkarni;Mitesh M. Khapra;Saurabh Sohoney;P. Bhattacharyya

文献摘要

被引文献

相似文献

我们描述了两种特定领域的全词词义消歧方法。第一种方法是一种基于知识的方法,提取特定领域的最大连接组件的Wordnet图,利用出现在特定领域的未标记语料库中的所有候选同义词集之间的语义关系。给定一个测试词,通过只考虑那些属于前k个最大连通分量的候选同义词集来执行消歧。 第二种方法是一种弱监督的方法,它依赖于“一个意义每域”的启发式,并使用一些手标记的例子中最频繁出现的单词在目标域。一旦最频繁的词已经被消除歧义,它们就可以使用迭代消除歧义算法为消除句子中的其他词的歧义提供强有力的线索。我们的弱监督系统在参与任务的所有系统中表现最好,即使它使用了来自目标域的100个手动标记的示例。
We describe two approaches for All-words Word Sense Disambiguation on a Specific Domain. The first approach is a knowledge based approach which extracts domain-specific largest connected components from the Wordnet graph by exploiting the semantic relations between all candidate synsets appearing in a domain-specific untagged corpus. Given a test word, disambiguation is performed by considering only those candidate synsets that belong to the top-k largest connected components. The second approach is a weakly supervised approach which relies on the "One Sense Per Domain" heuristic and uses a few hand labeled examples for the most frequently appearing words in the target domain. Once the most frequent words have been disambiguated they can provide strong clues for disambiguating other words in the sentence using an iterative disambiguation algorithm. Our weakly supervised system gave the best performance across all systems that participated in the task even when it used as few as 100 hand labeled examples from the target domain.