Efficient query answering in probabilistic RDF graphs

Efficient query answering in probabilistic RDF graphs
复制标题

DOI:
10.1145/1989323.1989341
复制
发表时间:
2011-06
期刊:
--
影响因子:
--
通讯作者:
Xiang Lian;Lei Chen
Xiang Lian;Lei Chen
中科院分区:
其他
文献类型:
--
作者:
Xiang Lian;Lei Chen

文献摘要

被引文献

相似文献

在本文中,我们解决了在概率RDF数据图上有效回答查询的问题。具体来说,我们通过概率图对RDF数据建模,并且RDF查询等价于对概率图的子图进行搜索,这些子图与给定的查询图有很高的匹配概率。为了有效地处理概率RDF图上的查询,我们提出了有效的剪枝机制,结构剪枝和概率剪枝。对于结构剪枝,我们考虑顶点/边缘标签的分布和其他结构信息,精心设计了剪枝概要,以提高剪枝能力。对于概率剪枝,我们推导了一个成本模型来指导概率上界的预计算,从而使查询成本预期较低。我们构建了一个集成了结构和概率剪枝的概要/统计的索引结构,并提出了一种有效的方法来回答对概率RDF图数据的查询。我们的解决方案的有效性已通过大量的实验得到验证。
In this paper, we tackle the problem of efficiently answering queries on probabilistic RDF data graphs. Specifically, we model RDF data by probabilistic graphs, and an RDF query is equivalent to a search over subgraphs of probabilistic graphs that have high probabilities to match with a given query graph. To efficiently processqueries on probabilistic RDF graphs, we propose effective pruning mechanisms, structural and probabilistic pruning. For the structural pruning, we carefully design synopses for vertex/edge labels by considering their distributions and other structural information, in order to improve the pruning power. For the probabilistic pruning, we derive a cost model to guide the pre-computation of probability upper bounds such that the query cost is expected to be low. We construct an index structure that integrates synopses/statistics for structural and robabilistic pruning, and propose an efficient approach to answer queries on probabilistic RDF graph data. The efficiency of our solutions has been verified through extensive experiments.