Discovering Relations Between Named Entities from a Large Raw Corpus Using Tree Similarity-Based Clustering

Discovering Relations Between Named Entities from a Large Raw Corpus Using Tree Similarity-Based Clustering
复制标题

DOI:
10.1007/11562214_34
复制
发表时间:
2005-10
期刊:
--
影响因子:
--
通讯作者:
Min Zhang;Jian Su;Danmei Wang;Guodong Zhou;C. Tan
Min Zhang;Jian Su;Danmei Wang;Guodong Zhou;C. Tan
中科院分区:
其他
文献类型:
--
作者:
Min Zhang;Jian Su;Danmei Wang;Guodong Zhou;C. Tan

文献摘要

被引文献

相似文献

我们提出了一种基于树相似性的无监督学习方法来从大型原始语料库中提取命名实体之间的关系。我们的方法将关系抽取作为一个聚类问题的浅解析树。首先,我们修改以前的树核关系提取,以更有效地估计解析树之间的相似性。然后,在层次聚类算法中使用解析树之间的相似性来将实体对分组到不同的聚类中。最后,每个聚类被标记的指示词和不可靠的聚类被修剪掉。对纽约时报(1995)语料的评价表明,我们的方法在F-测度上比以前的工作高5。它还表明,我们的方法在高频和低频实体对上都表现良好。据我们所知,这是第一个工作,使用树的相似性度量关系聚类。
We propose a tree-similarity-based unsupervised learning method to extract relations between Named Entities from a large raw corpus. Our method regards relation extraction as a clustering problem on shallow parse trees. First, we modify previous tree kernels on relation extraction to estimate the similarity between parse trees more efficiently. Then, the similarity between parse trees is used in a hierarchical clustering algorithm to group entity pairs into different clusters. Finally, each cluster is labeled by an indicative word and unreliable clusters are pruned out. Evaluation on the New York Times (1995) corpus shows that our method outperforms the only previous work by 5 in F-measure. It also shows that our method performs well on both high-frequent and less-frequent entity pairs. To the best of our knowledge, this is the first work to use a tree similarity metric in relation clustering.