SemStore: A Semantic-Preserving Distributed RDF Triple Store

SemStore: A Semantic-Preserving Distributed RDF Triple Store
复制标题

DOI:
10.1145/2661829.2661876
复制
发表时间:
2014-11
期刊:
Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Buwen Wu;Yongluan Zhou;Pingpeng Yuan;Hai Jin;Ling Liu
Buwen Wu;Yongluan Zhou;Pingpeng Yuan;Hai Jin;Ling Liu
中科院分区:
其他
文献类型:
--
作者:
Buwen Wu;Yongluan Zhou;Pingpeng Yuan;Hai Jin;Ling Liu

文献摘要

被引文献

相似文献

RDF数据模型的灵活性吸引了越来越多的组织以RDF格式存储其数据。随着RDF数据集的快速增长,为了提供理想的查询性能,部署计算节点集群来处理大规模RDF数据是不可避免的。在本文中,我们解决了横向扩展RDF引擎中的数据分区和查询优化的挑战性问题。我们发现,现有的方法只关注于使用细粒度的结构信息进行数据分区,因此无法定位许多类型的复杂查询。然后,我们提出了一种完全不同的方法,其中使用了一种粗粒度的结构,即有根子图(RSG)作为划分单元。通过这样做,我们可以在更大的范围内捕获结构信息,从而能够本地化许多复杂的查询。我们还提出了一种k-Means划分算法,用于将RSG分配到计算节点上,并提出了一种查询优化策略,以最小化查询处理过程中节点间的通信。使用基准数据集和真实数据集进行的广泛实验研究表明,我们的引擎SemStore在查询响应时间方面比现有系统高出几个数量级。
The flexibility of the RDF data model has attracted an increasing number of organizations to store their data in an RDF format. With the rapid growth of RDF datasets, we envision that it is inevitable to deploy a cluster of computing nodes to process large-scale RDF data in order to deliver desirable query performance. In this paper, we address the challenging problems of data partitioning and query optimization in a scale-out RDF engine. We identify that existing approaches only focus on using fine-grained structural information for data partitioning, and hence fail to localize many types of complex queries. We then propose a radically different approach, where a coarse-grained structure, namely Rooted Sub-Graph (RSG), is used as the partition unit. By doing so, we can capture structural information at a much greater scale and hence are able to localize many complex queries. We also propose a k-means partitioning algorithm for allocating the RSGs onto the computing nodes as well as a query optimization strategy to minimize the inter-node communication during query processing. An extensive experimental study using benchmark datasets and real dataset shows that our engine, SemStore, outperforms existing systems by orders of magnitudes in terms of query response time.