RuDiK: Rule Discovery in Knowledge Bases

RuDiK: Rule Discovery in Knowledge Bases
复制标题

RuDiK:知识库中的规则发现

DOI:
10.14778/3229863.3236231
复制
发表时间:
2018
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Paolo Papotti
Paolo Papotti
中科院分区:
--
文献类型:
--
作者:
Stefano Ortona;Venkata Vamsikrishna Meduri;Paolo Papotti

文献摘要

被引文献

相似文献

RuDiK是一个在知识库(KBs)上发现声明性规则的系统。鲁迪克发现了 积极 识别实体之间关系的规则,例如,“如果两个人有相同的父母,他们是兄弟姐妹”, 负 识别数据矛盾的规则,例如,“如果两个人结婚,其中一个不能是另一个的孩子”。规则可以帮助领域专家管理大型KB中的数据。积极规则建议新的事实,以减轻不完整性和消极规则检测错误的事实。此外,否定规则对于生成学习算法的否定示例是有用的。RuDiK超越了现有的解决方案,因为它发现规则, 表达规则语言 w.r.t.以前的方法,这导致知识库中的事实的广泛覆盖,并且它的挖掘对现有的 KB中的错误和不完整性。 该系统已部署在多个知识库中,包括Yago、DBpedia、Freebase和Wiki-Data,并分别以85%至97%的准确率识别新事实和真实的错误。本演示展示了如何使用RuDiK与领域专家进行交互。一旦观众选择了知识库和谓词,他们将添加新的事实,删除错误,并使用自动生成的示例训练机器学习系统。
RuDiK is a system for the discovery of declarative rules over knowledge-bases (KBs). RuDiK discovers both positive rules, which identify relationships between entities, e.g., "if two persons have the same parent, they are siblings", and negative rules, which identify data contradictions, e.g., "if two persons are married, one cannot be the child of the other". Rules help domain experts to curate data in large KBs. Positive rules suggest new facts to mitigate incompleteness and negative rules detect erroneous facts. Also, negative rules are useful to generate negative examples for learning algorithms. RuDiK goes beyond existing solutions since it discovers rules with a more expressive rule language w.r.t. previous approaches, which leads to wide coverage of the facts in the KB, and its mining is robust to existing errors and incompleteness in the KB. The system has been deployed for multiple KBs, including Yago, DBpedia, Freebase and Wiki-Data, and identifies new facts and real errors with 85% to 97% accuracy, respectively. This demonstration shows how RuDiK can be used to interact with domain experts. Once the audience pick a KB and a predicate, they will add new facts, remove errors, and train a machine learning system with automatically generated examples.