RDF Multi-query Optimization Algorithm for Query Rewriting Using Common Subgraphs

RDF Multi-query Optimization Algorithm for Query Rewriting Using Common Subgraphs
复制标题

DOI:
10.1145/3331453.3361278
复制
发表时间:
2019-10
期刊:
Proceedings of the 3rd International Conference on Computer Science and Application Engineering
影响因子:
--
通讯作者:
Manzi Wang;Haidong Fu;Fangfang Xu
Manzi Wang;Haidong Fu;Fangfang Xu
中科院分区:
其他
文献类型:
--
作者:
Manzi Wang;Haidong Fu;Fangfang Xu

文献摘要

被引文献

相似文献

在当前大数据快速发展的趋势下,RDF数据集上频繁出现高并发查询处理的应用场景。解决并发查询的多查询优化方案需要为由一组查询组成的查询集提供全局近似最优解,以最小化查询集的总体时间开销。在利用RDF存储索引加速统计和缩小语义剪枝范围的前提下,首先将简化的多个查询转换为连接图,然后对所有查询进行聚类和分组。在每个组中,迭代搜索连接图的所有公共子图,并建立映射表。然后,将公共子图按顶点数降序排列,构造查询重写方案。最后,对所有重写查询,采用基于选择率估计的动态规划算法进行二次优化。一方面,利用公共子图重写查询以减少查询次数,从而通过可重用的公共结果集来降低开销。另一方面,由于RDF存储索引的建立,可以快速估计选择率,并对改写后的查询进行再次优化,提高整体查询效率。实验结果表明,与现有的查询方案相比,该算法具有更好的查询性能,特别是在RDF数据集较大、查询集查询数量较多、查询语句复杂的情况下,本文提出的多查询优化方法效果更好。
In the current trend of rapid development of big data, there are frequent application scenarios of high concurrent query processing on RDF data sets. The multi-query optimization scheme for solving concurrent queries needs to provide a global approximate optimal solution for the query set composed of a set of queries, so as to minimize the overall time cost of the query set. Under the premise of accelerating statistics by RDF storage index and narrowing the scope of semantic pruning, firstly, simplified multiple queries are converted into connection graphs, and then all queries are clustered and grouped. In each group, all common subgraphs of connection graphs are iteratively searched and mapping tables are established. Then, the common subgraph is arranged in descending order by the number of vertices to construct the query rewriting scheme. Finally, for all the rewritten queries, the dynamic programming algorithm based on selection rate estimation is used for secondary optimization. On the one hand, the common subgraph is used to rewrite the query to reduce the number of queries, so as to reduce the cost through the reusable common result set. On the other hand, because of the establishment of RDF storage index, the selection rate can be estimated quickly, and the rewritten queries can be optimized again to improve the overall query efficiency. The experimental results show that the proposed algorithm has better query performance than the existing query schemes, especially when the RDF dataset is large, the number of queries in the query set is large, and the query statements are complex, the multi-query optimization method in this paper works better.