Scalable Processing of Contemporary Semi-Structured Data on Commodity Parallel Processors - A Compilation-based Approach
Scalable Processing of Contemporary Semi-Structured Data on Commodity Parallel Processors - A Compilation-based Approach
复制标题
商品并行处理器上当代半结构化数据的可扩展处理 - 基于编译的方法
DOI:
10.1145/3297858.3304008
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Zhao, Zhijia
中科院分区:
文献类型:
--
作者:
Jiang, Lin;Sun, Xiaofan;Farooq, Umar;Zhao, Zhijia
JSON (JavaScript Object Notation) and its derivatives are essential in the modern computing infrastructure. However, existing software often fails to process such types of data in a scalable way, mainly for two reasons: (i) the processing often requires to build a memory-consuming parse tree; (ii) there exist inherent dependences in processing the data stream, preventing any data-level parallelization. Facing the challenges, developers often have to construct ad-hoc pre-parsers to split the data stream in order to reduce the memory consumption and increase the data parallelism. However, this strategy requires more programming efforts. Moreover, the pre-parsing itself is non-trivial to parallelize, thus introducing a new serial bottleneck. To solve the dilemma, this work introduces a scalable yet fully automatic solution - a compilation system, namely JPStream, that compiles standard JSONPath queries into parallel executables with bounded memory footprints. First, JPStream adopts a stream processing design that combines the querying and parsing into one pass, without generating any in-memory parse tree. To achieve this, JPStream uses a novel joint compilation technique that compiles the queries and the JSON syntax together into a single automaton. Furthermore, JPStream leverages the "enumerability'' of automaton to break the dependences and reason about the transition rules to prune infeasible states. It also features a runtime that learns structural constraints from the input to enhance the pruning. Evaluation on real-world JSON datasets with standard JSONPath queries shows that JPStream can reduce the memory consumption significantly, by up to 95%, meanwhile achieving near-linear speedup on multicore and manycore processors.
登录
查看更多内容
DOI:
10.1109/micro.2018.00079
发表时间:
2018
期刊:
2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO
影响因子:
--
作者:
Angstadt, Kevin;Subramaniyan, Arun;Sadredini, Elaheh;Rahimi, Reza;Skadron, Kevin;Weimer, Westley;Das, Reetuparna
通讯作者:
Das, Reetuparna
DOI:
--
发表时间:
2016
期刊:
SIGMOD Conference
影响因子:
--
作者:
Peter Ogden;David B. Thomas;P. Pietzuch
通讯作者:
P. Pietzuch
DOI:
10.1145/3079079.3079082
发表时间:
2017
期刊:
Proceedings of the International Conference on Supercomputing
影响因子:
--
作者:
Junqiao Qiu;Zhijia Zhao;Bo Wu;Abhinav Vishnu;S. Song
通讯作者:
S. Song
DOI:
--
发表时间:
2018
期刊:
International Conference on Extending Database Technology
影响因子:
--
作者:
C. Pavlopoulou;E. Carman;T. Westmann;M. Carey;V. Tsotras
通讯作者:
V. Tsotras
DOI:
10.1145/2967938.2967965
发表时间:
2016
期刊:
2016 International Conference on Parallel Architecture and Compilation Techniques (PACT)
影响因子:
--
作者:
Junqiao Qiu;Zhijia Zhao;Bin Ren
通讯作者:
Bin Ren