An assertion and alignment correction framework for large scale knowledge bases

An assertion and alignment correction framework for large scale knowledge bases
复制标题

DOI:
10.3233/sw-210448
复制
发表时间:
2021-10
期刊:
影响因子:
3
通讯作者:
Jiaoyan Chen;E. Jiménez-Ruiz;Ian Horrocks;Xi Chen;E. B. Myklebust
Jiaoyan Chen;E. Jiménez-Ruiz;Ian Horrocks;Xi Chen;E. B. Myklebust
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jiaoyan Chen;E. Jiménez-Ruiz;Ian Horrocks;Xi Chen;E. B. Myklebust

文献摘要

相似文献

通过从百科全书、文本和表格中提取信息,以及对多个来源进行对齐,构建了各种知识库(KBs)。它们的有用性和可用性常常受到质量问题的限制。一个常见的问题是存在错误的断言和对齐,这通常是由词汇或语义混淆引起的。研究了此类断言和对齐的纠错问题,提出了一种集词法匹配、上下文感知sub-KB提取、语义嵌入、软约束挖掘和语义一致性检查于一体的通用纠错框架。该框架使用来自DBpedia的一组文字断言、来自企业医疗知识库的一组实体断言和来自通过集成Wikidata、Discogs和MusicBrainz构建的音乐知识库的一组映射断言来评估。它取得了令人满意的结果,正确率(即目标断言/比对被正确替换的比例)分别为70.1%、60.9%和71.8%。
Various knowledge bases (KBs) have been constructed via information extraction from encyclopedias, text and tables, as well as alignment of multiple sources. Their usefulness and usability is often limited by quality issues. One common issue is the presence of erroneous assertions and alignments, often caused by lexical or semantic confusion. We study the problem of correcting such assertions and alignments, and present a general correction framework which combines lexical matching, context-aware sub-KB extraction, semantic embedding, soft constraint mining and semantic consistency checking. The framework is evaluated with one set of literal assertions from DBpedia, one set of entity assertions from an enterprise medical KB, and one set of mapping assertions from a music KB constructed by integrating Wikidata, Discogs and MusicBrainz. It has achieved promising results, with a correction rate (i.e., the ratio of the target assertions/alignments that are corrected with right substitutes) of 70.1 %, 60.9 % and 71.8 %, respectively.