Second Order Co-occurrence PMI for Determining the Semantic Similarity of Words

Second Order Co-occurrence PMI for Determining the Semantic Similarity of Words
复制标题

DOI:
--
复制
发表时间:
2006-05
期刊:
--
影响因子:
--
通讯作者:
Aminul Islam;Diana Inkpen
Aminul Islam;Diana Inkpen
中科院分区:
其他
文献类型:
--
作者:
Aminul Islam;Diana Inkpen

文献摘要

被引文献

相似文献

本文提出了一种基于语料库的新方法来计算两个目标词的语义相似度。我们的方法称为二阶共现PMI (SOC-PMI),使用逐点互信息对两个目标单词的重要相邻单词列表进行排序。然后我们考虑两个列表中常见的单词并聚合它们的 PMI 值(来自相反的列表)以计算相对语义相似度。我们的方法使用 Miller 和 Charler (1991) 的 30 个名词对子集、Ruben-stein 和 Goodenough’s (1965) 的 65 个名词对、来自英语作为外语测试 (TOEFL) 的 80 个同义词测试问题以及来自英语作为第二语言 (ESL) 测试集合的 50 个同义词测试问题进行了实证评估。评估结果表明,我们的方法优于几种基于语料库的竞争方法。
This paper presents a new corpus-based method for calculating the semantic similarity of two target words. Our method, called Second Order Co-occurrencePMI (SOC-PMI), uses Pointwise Mutual Information to sort lists of important neighbor words of the two target words. Then we consider the words which are common in both lists and aggregate their PMI values (from the opposite list) to calculate the relative semantic similarity. Our method was empirically evaluated using Miller and Charler’s (1991) 30 noun pair subset, Ruben-stein and Goodenough’s (1965) 65 noun pairs, 80 synonym test questions from the Test of English as a Foreign Language (TOEFL), and 50 synonym test questions from a collection of English as a Second Language (ESL) tests. Evaluation results show that our method outperforms several competing corpus-based methods.