A self-training approach for resolving object coreference on the semantic web

A self-training approach for resolving object coreference on the semantic web
复制标题

DOI:
10.1145/1963405.1963421
复制
发表时间:
2011-03
期刊:
--
影响因子:
--
通讯作者:
Wei Hu;Jianfeng Chen;Yuzhong Qu
Wei Hu;Jianfeng Chen;Yuzhong Qu
中科院分区:
其他
文献类型:
--
作者:
Wei Hu;Jianfeng Chen;Yuzhong Qu

文献摘要

被引文献

相似文献

语义Web上的对象可能由不同的方用多个URI表示。对象共指解析是识别表示相同对象的“等价”URI。在链接开放数据(LOD)倡议的推动下,数百万个URI已经显式地与owl:sameAs语句相关联,但潜在的共指关系仍然相当可观。现有的方法主要从两个方向来解决这个问题:一个是基于OWL语义授权的等价推理,它发现语义相关的URI,但可能会忽略许多潜在的;另一个是通过属性值对之间的相似性计算,这并不总是足够准确。在本文中,我们提出了一个自我训练的方法,在语义Web上的对象共指解析,利用这两类方法之间的差距,语义共指URI和潜在的候选人的桥梁差距。对于一个对象URI,我们首先建立了一个内核,它由语义相关的URI基于owl:sameAs,(逆)功能属性和(最大)基数,然后迭代扩展这样的内核的判别属性值对的URI的描述。特别是,可辨别性是学习与统计测量,它不仅利用关键特征代表一个对象,但也考虑到从语用属性之间的匹配性。此外,频繁的属性组合挖掘,以提高分辨率的准确性。我们实现了一个可扩展的系统,并证明了我们的方法实现了良好的精度和召回解决对象共指,在基准和大规模的数据集。
An object on the Semantic Web is likely to be denoted with multiple URIs by different parties. Object coreference resolution is to identify "equivalent" URIs that denote the same object. Driven by the Linking Open Data (LOD) initiative, millions of URIs have been explicitly linked with owl:sameAs statements, but potentially coreferent ones are still considerable. Existing approaches address the problem mainly from two directions: one is based upon equivalence inference mandated by OWL semantics, which finds semantically coreferent URIs but probably omits many potential ones; the other is via similarity computation between property-value pairs, which is not always accurate enough. In this paper, we propose a self-training approach for object coreference resolution on the Semantic Web, which leverages the two classes of approaches to bridge the gap between semantically coreferent URIs and potential candidates. For an object URI, we firstly establish a kernel that consists of semantically coreferent URIs based on owl:sameAs, (inverse) functional properties and (max-)cardinalities, and then extend such kernel iteratively in terms of discriminative property-value pairs in the descriptions of URIs. In particular, the discriminability is learnt with a statistical measurement, which not only exploits key characteristics for representing an object, but also takes into account the matchability between properties from pragmatics. In addition, frequent property combinations are mined to improve the accuracy of the resolution. We implement a scalable system and demonstrate that our approach achieves good precision and recall for resolving object coreference, on both benchmark and large-scale datasets.