RDF2Vec: RDF graph embeddings and their applications

RDF2Vec: RDF graph embeddings and their applications
复制标题

DOI:
10.3233/sw-180317
复制
发表时间:
2019-01-01
期刊:
影响因子:
3
通讯作者:
Paulheim, Heiko
Paulheim, Heiko
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ristoski, Petar;Rosati, Jessica;Paulheim, Heiko

文献摘要

被引文献

相似文献

在许多数据挖掘和信息检索任务中,关联开放数据被认为是一种有价值的背景信息来源。然而,大多数现有工具需要命题形式的特征,即,与实例相关联的名义或数值特征的向量,而链接开放数据源本质上是图形。在本文中,我们提出了RDF 2 Vec,一种使用语言建模方法从单词序列中进行无监督特征提取的方法,并使其适应RDF图。我们通过利用Weisfeiler-Lehman子树RDF图内核和图遍历收获的图子结构的本地信息来生成序列,并学习RDF图中实体的潜在数值表示。我们在三个不同的任务上评估我们的方法:(i)标准机器学习任务,(ii)实体和文档建模,以及(iii)基于内容的推荐系统。评估结果表明,所提出的实体嵌入优于现有技术,并且DBpedia和Wikidata等通用知识图的预先计算的特征向量表示可以很容易地重复用于不同的任务。
Linked Open Data has been recognized as a valuable source for background information in many data mining and information retrieval tasks. However, most of the existing tools require features in propositional form, i.e., a vector of nominal or numerical features associated with an instance, while Linked Open Data sources are graphs by nature. In this paper, we present RDF2Vec, an approach that uses language modeling approaches for unsupervised feature extraction from sequences of words, and adapts them to RDF graphs. We generate sequences by leveraging local information from graph sub-structures, harvested by Weisfeiler-Lehman Subtree RDF Graph Kernels and graph walks, and learn latent numerical representations of entities in RDF graphs. We evaluate our approach on three different tasks: (i) standard machine learning tasks, (ii) entity and document modeling, and (iii) content-based recommender systems. The evaluation shows that the proposed entity embeddings outperform existing techniques, and that pre-computed feature vector representations of general knowledge graphs such as DBpedia and Wikidata can be easily reused for different tasks.