Substructure counting graph kernels for machine learning from RDF data

Substructure counting graph kernels for machine learning from RDF data
复制标题

DOI:
10.1016/j.websem.2015.08.002
复制
发表时间:
2015-12-01
影响因子:
2.5
通讯作者:
de Rooij, Steven
de Rooij, Steven
中科院分区:
计算机科学2区
文献类型:
--
作者:
de Vries, Gerben Klaas Dirk;de Rooij, Steven

文献摘要

被引文献

相似文献

在本文中,我们介绍了一个框架,用于学习RDF数据使用图内核,计数RDF图中的子结构,系统地涵盖了大多数现有的内核以前定义的,并提供了一些新的变种。我们的定义包括直接在RDF图上计算的快速内核变体。为了提高这些内核的性能,我们详细介绍了两种策略。第一个策略涉及忽略在实例中具有低频率的顶点标签。我们的第二个策略是删除中心节点以简化RDF图。我们用真实世界的RDF数据集在一些分类实验中测试了我们的内核。总的来说,计算子树的内核表现出最好的性能。然而,紧随其后的是简单的标签袋基线内核。直接内核大大减少了计算时间,同时保持性能不变。对于步行计数内核,计算时间的减少如此之大,从而成为计算上可行的内核。忽略低频标签可以提高所有数据集的性能。中心删除算法提高了我们三分之二的较小数据集的性能,但在我们的较大数据集上使用时几乎没有影响。(C)2015爱思唯尔B.V.保留所有权利。
In this paper we introduce a framework for learning from RDF data using graph kernels that count substructures in RDF graphs, which systematically covers most of the existing kernels previously defined and provides a number of new variants. Our definitions include fast kernel variants that are computed directly on the RDF graph. To improve the performance of these kernels we detail two strategies. The first strategy involves ignoring the vertex labels that have a low frequency among the instances. Our second strategy is to remove hubs to simplify the RDF graphs. We test our kernels in a number of classification experiments with real-world RDF datasets. Overall the kernels that count subtrees show the best performance. However, they are closely followed by simple bag of labels baseline kernels. The direct kernels substantially decrease computation time, while keeping performance the same. For the walks counting kernel this decrease in computation time is so large that it thereby becomes a computationally viable kernel to use. Ignoring low frequency labels improves the performance for all datasets. The hub removal algorithm increases performance on two out of three of our smaller datasets, but has little impact when used on our larger datasets. (C) 2015 Elsevier B.V. All rights reserved.