Factorizing YAGO: scalable machine learning for linked data

Factorizing YAGO: scalable machine learning for linked data
复制标题

DOI:
10.1145/2187836.2187874
复制
发表时间:
2012-04
期刊:
Proceedings of the 21st international conference on World Wide Web
影响因子:
--
通讯作者:
Maximilian Nickel;Volker Tresp;H. Kriegel
Maximilian Nickel;Volker Tresp;H. Kriegel
中科院分区:
其他
文献类型:
--
作者:
Maximilian Nickel;Volker Tresp;H. Kriegel

文献摘要

被引文献

相似文献

海量的结构化信息已经在语义网的链接开放数据(LOD)云中发布,并且它们的规模仍在快速增长。然而,由于LOD的大小、部分数据的不一致和固有的噪声,通过推理和查询来获取这些信息有时是困难的。机器学习提供了一种利用LOD数据的替代方法,其优点是机器学习算法通常对噪声和数据不一致都具有健壮性,并且能够有效地利用数据中的非确定性依赖关系。从机器学习的角度来看,由于LOD的关系性质和规模,LOD具有挑战性。在这里,我们提出了一种在LOD数据上进行关系学习的有效方法,该方法基于稀疏张量的因式分解,该稀疏张量可以扩展到由数百万个实体、数百个关系和数十亿个已知事实组成的数据。此外,我们还展示了如何将本体知识整合到因式分解中以改善学习结果,以及如何将计算分布在多个节点上。我们证明了我们的方法能够对Yago~2核心本体进行分解,并使用一台双核桌面计算机对这个庞大的知识库进行全局预测。此外,我们的实验表明,我们的方法在几个与关联数据相关的关系学习任务中取得了良好的结果。一旦计算出因式分解,我们的模型就能够有效地预测YAGO~2核心本体中4.3YAG1014可能的三元组中的任何一个的可能性,而不需要任何额外的训练。
Vast amounts of structured information have been published in the Semantic Web's Linked Open Data (LOD) cloud and their size is still growing rapidly. Yet, access to this information via reasoning and querying is sometimes difficult, due to LOD's size, partial data inconsistencies and inherent noisiness. Machine Learning offers an alternative approach to exploiting LOD's data with the advantages that Machine Learning algorithms are typically robust to both noise and data inconsistencies and are able to efficiently utilize non-deterministic dependencies in the data. From a Machine Learning point of view, LOD is challenging due to its relational nature and its scale. Here, we present an efficient approach to relational learning on LOD data, based on the factorization of a sparse tensor that scales to data consisting of millions of entities, hundreds of relations and billions of known facts. Furthermore, we show how ontological knowledge can be incorporated in the factorization to improve learning results and how computation can be distributed across multiple nodes. We demonstrate that our approach is able to factorize the YAGO~2 core ontology and globally predict statements for this large knowledge base using a single dual-core desktop computer. Furthermore, we show experimentally that our approach achieves good results in several relational learning tasks that are relevant to Linked Data. Once a factorization has been computed, our model is able to predict efficiently, and without any additional training, the likelihood of any of the 4.3 ⋅ 1014 possible triples in the YAGO~2 core ontology.