A Parallel and Scalable Processor for JSON Data

A Parallel and Scalable Processor for JSON Data
复制标题

JSON 数据的并行且可扩展的处理器

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Extending Database Technology
影响因子:
--
通讯作者:
V. Tsotras
V. Tsotras
中科院分区:
--
文献类型:
--
作者:
C. Pavlopoulou;E. Carman;T. Westmann;M. Carey;V. Tsotras

文献摘要

被引文献

相似文献

对JSON数据日益增长的兴趣产生了对其高效处理的需求。尽管JSON是一种简单的数据交换格式,但它的查询并不总是有效的,特别是在大型数据存储库的情况下。这项工作旨在将XQuery语言规范的JSONiq扩展集成到现有的查询处理器(Apache VXQuery)中,使其能够并行查询JSON数据。VXQuery建立在hyrack(一个生成并行作业的框架)和Algebricks(一个与语言无关的查询代数工具箱)之上,可以动态地处理数据,这与其他需要首先加载数据的知名系统形成了鲜明对比。因此,消除了数据加载的额外成本。在本文中,我们实现了三类重写规则,这些规则利用了上述平台的特征来有效地处理路径表达式,并引入了查询内并行性。我们使用一个大的(803GB)传感器读数数据集来评估我们的实现。我们的结果表明,所提出的重写规则导致JSON数据的高效和可扩展的并行处理。
Increasing interest in JSON data has created a need for its efficient processing. Although JSON is a simple data exchange format, its querying is not always effective, especially in the case of large repositories of data. This work aims to integrate the JSONiq extension to the XQuery language specification into an existing query processor (Apache VXQuery) to enable it to query JSON data in parallel. VXQuery is built on top of Hyracks (a framework that generates parallel jobs) and Algebricks (a language-agnostic query algebra toolbox) and can process data on the fly, in contrast to other well-known systems which need to load data first. Thus, the extra cost of data loading is eliminated. In this paper, we implement three categories of rewrite rules which exploit the features of the above platforms to efficiently handle path expressions along with introducing intra-query parallelism. We evaluate our implementation using a large (803GB) dataset of sensor readings. Our results show that the proposed rewrite rules lead to efficient and scalable parallel processing of JSON data.