Infrastructure for Efficient Exploration of Large Scale Linked Data via Contextual Tag Clouds

Infrastructure for Efficient Exploration of Large Scale Linked Data via Contextual Tag Clouds
复制标题

通过上下文标签云高效探索大规模关联数据的基础设施

DOI:
--
复制
发表时间:
2013
期刊:
International Workshop on the Semantic Web
影响因子:
--
通讯作者:
J. Heflin
J. Heflin
中科院分区:
--
文献类型:
--
作者:
Xingjian Zhang;Dezhao Song;S. Priya;J. Heflin

文献摘要

被引文献

相似文献

在本文中,我们提出了上下文标签云系统的基础架构,该系统可以执行大量关于使用特定本体术语的实例数量的查询。上下文标记云系统是一个帮助用户探索大规模RDF数据集的新颖应用程序:标记是本体术语(类和属性),上下文是定义实例子集的一组标记,字体大小反映使用每个标记的实例的数量。它可视化由用户构造的上下文指定的实例模式。给定具有特定上下文的请求,系统需要快速查找上下文中的实例使用的其他标记,以及上下文中有多少实例使用每个标记。我们在本文中回答的关键问题是如何扩展到关联数据;特别地,我们使用了一个包含14亿个三元组和超过38万个标签的数据集。当由用户指导时,计算应该考虑本体中分类法和/或域/范围公理的蕴涵,这一事实使情况变得复杂。我们将可扩展的预处理方法与特殊构造的倒排索引相结合,并使用三种方法来修剪不必要的计数,以实现更快的交叉计算。我们将我们的系统与最先进的三重存储进行比较,检查修剪规则如何与推理相互作用,并分析我们的设计选择。
In this paper we present the infrastructure of the contextual tag cloud system which can execute large volumes of queries about the number of instances that use particular ontological terms. The contextual tag cloud system is a novel application that helps users explore a large scale RDF dataset: the tags are ontological terms (classes and properties), the context is a set of tags that defines a subset of instances, and the font sizes reflect the number of instances that use each tag. It visualizes the patterns of instances specified by the context a user constructs. Given a request with a specific context, the system needs to quickly find what other tags the instances in the context use, and how many instances in the context use each tag. The key question we answer in this paper is how to scale to Linked Data; in particular we use a dataset with 1.4 billion triples and over 380,000 tags. This is complicated by the fact that the calculation should, when directed by the user, consider the entailment of taxonomic and/or domain/range axioms in the ontology. We combine a scalable preprocessing approach with a specially-constructed inverted index and use three approaches to prune unnecessary counts for faster intersection computations. We compare our system with a state-of-the-art triple store, examine how pruning rules interact with inference and analyze our design choices.