Querying Web-Scale Knowledge Graphs Through Effective Pruning of Search Space
Querying Web-Scale Knowledge Graphs Through Effective Pruning of Search Space
复制标题
通过有效修剪搜索空间查询网络规模的知识图
DOI:
10.1109/tpds.2017.2665478
复制
发表时间:
2017-08-01
影响因子:
5.3
通讯作者:
Gao, Lixin
中科院分区:
文献类型:
--
作者:
Jin, Jiahui;Luo, Junzhou;Gao, Lixin
Web-scale knowledge graphs containing billions of entities are common nowadays. Querying these graphs can be modeled as a subgraph matching problem. Since knowledge graphs are incomplete and noisy in nature, it is important to discover answers matching exactly as well as answers similar to queries. Existing graph matching algorithms usually use graph indices to accelerate query processing. For billion-node graphs, it may be infeasible to build the graph indices due to the amount of work and the memory/storage required. In this paper, we propose an efficient algorithm for finding the best <inline-formula><tex-math notation="LaTeX">$k$</tex-math><alternatives> <inline-graphic xlink:href="jin-ieq1-2665478.gif"/></alternatives></inline-formula> answers for a given query without precomputing graph indices. An answer’s quality is measured by a matching score that is computed online. To accelerate query processing, we propose a novel technique for bounding the matching scores during the computation. By using bounds, the low quality answers can be efficiently pruned. The bounding technique can be implemented in a distributed environment, allowing our approach to efficiently query web-scale knowledge graphs. We evaluate the effectiveness and the efficiency of our approach on real-world datasets. The result shows that our bounding technique can reduce the running time up to two orders of magnitude comparing to an approach without using bounds.