Sempala: Interactive SPARQL Query Processing on Hadoop

Sempala: Interactive SPARQL Query Processing on Hadoop
复制标题

DOI:
10.1007/978-3-319-11964-9_11
复制
发表时间:
2014-10
期刊:
--
影响因子:
--
通讯作者:
A. Schätzle;Martin Przyjaciel-Zablocki;Antony Neu;G. Lausen
A. Schätzle;Martin Przyjaciel-Zablocki;Antony Neu;G. Lausen
中科院分区:
其他
文献类型:
--
作者:
A. Schätzle;Martin Przyjaciel-Zablocki;Antony Neu;G. Lausen

文献摘要

被引文献

相似文献

在Schema.org等项目的推动下,语义注释数据的数量预计将稳步增长到大规模,需要基于集群的解决方案来查询它。与此同时,Hadoop已经成为大数据处理领域的主导,大型基础设施已经部署并用于多个应用领域。对于基于Hadoop的应用程序,公共数据池(HDFS)提供了许多协同优势,因此使用这些基础设施进行语义数据处理也非常有吸引力。事实上,现有的SPARQL-on- Hadoop(MapReduce)方法已经展示了非常好的可扩展性,然而,由于底层批处理框架,查询运行时相当慢。虽然这对于数据密集型查询是可以接受的,但对于大多数SPARQL查询来说,这并不令人满意,因为SPARQL查询通常更具选择性,只需要很小的数据子集。在本文中,我们介绍了Sempala,这是一种基于Hadoop的SPARQL over SQL-on-Hadoop方法,设计时考虑了选择性查询。我们的评估显示,与现有方法相比,性能提高了一个数量级,为Hadoop上的交互式SPARQL查询处理铺平了道路。
Driven by initiatives like Schema.org, the amount of semantically annotated data is expected to grow steadily towards massive scale, requiring cluster-based solutions to query it. At the same time, Hadoop has become dominant in the area of Big Data processing with large infrastructures being already deployed and used in manifold application fields. For Hadoop-based applications, a common data pool (HDFS) provides many synergy benefits, making it very attractive to use these infrastructures for semantic data processing as well. Indeed, existing SPARQL-on- Hadoop (MapReduce) approaches have already demonstrated very good scalability, however, query runtimes are rather slow due to the underlying batch processing framework. While this is acceptable for data-intensive queries, it is not satisfactory for the majority of SPARQL queries that are typically much more selective requiring only small subsets of the data. In this paper, we presentSempala, a SPARQL-over-SQL-on-Hadoop approach designed with selective queries in mind. Our evaluation shows performance improvements by an order of magnitude compared to existing approaches, paving the way for interactive-time SPARQL query processing on Hadoop.