Compressed Indexes for Fast Search of Semantic Data

Compressed Indexes for Fast Search of Semantic Data
复制标题

DOI:
10.1109/tkde.2020.2966609
复制
发表时间:
2021-09-01
影响因子:
8.9
通讯作者:
Venturini, Rossano
Venturini, Rossano
中科院分区:
计算机科学2区
文献类型:
--
作者:
Perego, Raffaele;Pibiri, Giulio Ermanno;Venturini, Rossano

文献摘要

被引文献

相似文献

RDF数据量的急剧增加需要有效的解决方案的三元组索引的问题,即设计一个压缩的数据结构,以压缩表示RDF三元组,通过保证,在同一时间,快速模式匹配操作。这个问题的核心在于为大型RDF数据集上的复杂SPARQL查询提供良好的实际性能。在这项工作中,我们提出了一个基于trie的索引布局来解决这个问题,并引入了两种新的技术,以减少其空间的表示,以提高效率。在广泛的公开可用的真实世界数据集上进行的广泛的实验分析表明,我们最好的空间/时间权衡配置大大优于现有的解决方案,通过减少30- 60%的空间和加快查询执行2 - 81倍。
The sheer increase in volume of RDF data demands efficient solutions for the triple indexing problem, that is to devise a compressed data structure to compactly represent RDF triples by guaranteeing, at the same time, fast pattern matching operations. This problem lies at the heart of delivering good practical performance for the resolution of complex SPARQL queries on large RDF datasets. In this work, we propose a trie-based index layout to solve the problem and introduce two novel techniques to reduce its space of representation for improved effectiveness. The extensive experimental analysis, conducted over a wide range of publicly available real-world datasets, reveals that our best space/time trade-off configuration substantially outperforms existing solutions at the state-of-the-art, by taking 30-60 percent less space and speeding up query execution by a factor of 2 - 81x.