Statistics-based parallelization of XPath queries in shared memory systems

Statistics-based parallelization of XPath queries in shared memory systems
复制标题

DOI:
10.1145/1739041.1739063
复制
发表时间:
2010-03
期刊:
--
影响因子:
--
通讯作者:
R. Bordawekar;Lipyeow Lim;Anastasios Kementsietsidis;Bryant Wei-Lun Kok
R. Bordawekar;Lipyeow Lim;Anastasios Kementsietsidis;Bryant Wei-Lun Kok
中科院分区:
其他
文献类型:
--
作者:
R. Bordawekar;Lipyeow Lim;Anastasios Kementsietsidis;Bryant Wei-Lun Kok

文献摘要

被引文献

相似文献

商品多核系统的广泛可用性为解决XML查询处理的延迟问题提供了机会。但是,仅在多个内核上执行多个XML查询,就可以解决吞吐量问题:需要进行拼写并行化即可利用多个处理核心以提高延迟。为了这项努力,本文调查了各个X Path的查询对共享 - 地址空间多核处理器的并行化。在分布式设置中并行化XPATH的许多以前的工作都无法利用多核系统的共享存储器并行性。我们提出了一个新颖的端到端并行化框架,该框架决定了并行化XML查询的最佳方法。该决定基于一种基于统计的方法,该方法既依赖查询细节和数据统计信息。在平行化过程的每个阶段,我们评估了三种替代方法,即数据,查询和混合分区。对于给定的XPATH查询,我们的并行化算法使用XML统计数据来估计这些不同替代方案的相对效率,并找到一个最佳的并行XPATH处理计划。我们使用知名XML文档的实验验证了我们的并行成本模型和优化框架,并证明可以使用商品多核系统加速XPATH处理。
The wide availability of commodity multi-core systems presents an opportunity to address the latency issues that have plaqued XML query processing. However, simply executing multiple XML queries over multiple cores merely addresses the throughput issue: intra-query parallelization is needed to exploit multiple processing cores for better latency. Toward this effort, this paper investigates the parallelization of individual XPath queries over shared-address space multi-core processors. Much previous work on parallelizing XPath in a distributed setting failed to exploit the shared memory parallelism of multi-core systems. We propose a novel, end-to-end parallelization framework that determines the optimal way of parallelizing an XML query. This decision is based on a statistics-based approach that relies both on the query specifics and the data statistics. At each stage of the parallelization process, we evaluate three alternative approaches, namely, data-, query-, and hybrid-partitioning. For a given XPath query, our parallelization algorithm uses XML statistics to estimate the relative efficiencies of these different alternatives and find an optimal parallel XPath processing plan. Our experiments using well-known XML documents validate our parallel cost model and optimization framework, and demonstrate that it is possible to accelerate XPath processing using commodity multi-core systems.