Detecting Errors in Numerical Linked Data Using Cross-Checked Outlier Detection

Detecting Errors in Numerical Linked Data Using Cross-Checked Outlier Detection
复制标题

使用交叉检查异常值检测来检测数字关联数据中的错误

DOI:
10.1007/978-3-319-11964-9_23
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Christian Bizer
Christian Bizer
中科院分区:
--
文献类型:
--
作者:
Daniel Fleischhacker;Heiko Paulheim;Volha Bryl;Johanna Völker;Christian Bizer

文献摘要

参考文献

被引文献

相似文献

用于识别数据中错误值的离群值检测通常应用于单个数据集,以搜索它们的非预期行为值。在这项工作中,我们提出了一种方法,它结合了两个独立的离群值检测运行的结果,以获得更可靠的结果,并防止自然离群值所产生的问题,这些离群值是数据集中的异常值,但仍然是正确的。关联数据特别适合于这种思想的应用,因为它提供了大量富含层次信息的数据,并且还包含实例之间的显式链接。在第一步中,我们将离群值检测方法应用于从单个存储库中提取的属性值,使用一种新的方法将数据拆分为相关子集。对于第二步,我们利用实例的owl:sameAs链接来获取额外的属性值,并对这些值执行第二次离群值检测。这样做可以让我们确认或拒绝错误值的评估。在DBpedia和NELL数据集上的实验证明了该方法的可行性。
Outlier detection used for identifying wrong values in data is typically applied to single datasets to search them for values of unexpected behavior. In this work, we instead propose an approach which combines the outcomes of two independent outlier detection runs to get a more reliable result and to also prevent problems arising from natural outliers which are exceptional values in the dataset but nevertheless correct. Linked Data is especially suited for the application of such an idea, since it provides large amounts of data enriched with hierarchical information and also contains explicit links between instances. In a first step, we apply outlier detection methods to the property values extracted from a single repository, using a novel approach for splitting the data into relevant subsets. For the second step, we exploit owl:sameAs links for the instances to get additional property values and perform a second outlier detection on these values. Doing so allows us to confirm or reject the assessment of a wrong value. Experiments on the DBpedia and NELL datasets demonstrate the feasibility of our approach.
DOI: 10.1007/978-3-642-35173-0
发表时间: 2012
期刊: 2012 IEEE Sixth International Conference on Semantic Computing
影响因子: --
作者:
P. Cudré-Mauroux;J. Heflin;E. Sirin;Tania Tudorache;J. Euzenat;M. Hauswirth;J. Parreira;J. Hendler;G. Schreiber;A. Bernstein;E. Blomqvist
通讯作者: P. Cudré-Mauroux;J. Heflin;E. Sirin;Tania Tudorache;J. Euzenat;M. Hauswirth;J. Parreira;J. Hendler;G. Schreiber;A. Bernstein;E. Blomqvist
具有数值属性的基于相关性的规则细化
DOI: --
发表时间: 2014
期刊: The Florida AI Research Society
影响因子: --
作者:
André Melo;M. Theobald;Johanna Völker
通讯作者: Johanna Völker
Nell2RDF:阅读 Web,并将其转换为 RDF
DOI: --
发表时间: 2013
期刊: KNOW@LOD
影响因子: --
作者:
Antoine Zimmermann;C. Gravier;Julien Subercaze;Quentin Cruzille
通讯作者: Quentin Cruzille
学习跨语言维基百科数据融合的冲突解决策略
DOI: --
发表时间: 2014
期刊: The Web Conference
影响因子: --
作者:
Volha Bryl;Christian Bizer
通讯作者: Christian Bizer
本体匹配,第二版
DOI: --
发表时间: 2013
期刊:
影响因子: --
作者:
J. Euzenat;P. Shvaiko
通讯作者: P. Shvaiko