Morsel-driven parallelism: a NUMA-aware query evaluation framework for the many-core age

Morsel-driven parallelism: a NUMA-aware query evaluation framework for the many-core age
复制标题

DOI:
10.1145/2588555.2610507
复制
发表时间:
2014-06
期刊:
Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data
影响因子:
--
通讯作者:
Viktor Leis;P. Boncz;A. Kemper;Thomas Neumann
Viktor Leis;P. Boncz;A. Kemper;Thomas Neumann
中科院分区:
其他
文献类型:
--
作者:
Viktor Leis;P. Boncz;A. Kemper;Thomas Neumann

文献摘要

被引文献

相似文献

随着现代计算机架构的发展,并行查询执行中最先进的方法面临着两个问题:(i)为了利用多核,所有查询工作必须在(很快)数百个线程之间均匀分布,以实现良好的加速,但(ii)由于现代乱序核心的复杂性,即使有准确的数据统计,均匀划分工作也很困难。因此,现有的计划驱动并行方法遇到了负载平衡和上下文切换瓶颈,因此不再可扩展。众核架构面临的第三个问题是内存控制器的分散化,这导致了非统一内存访问(NUMA)。作为回应,我们提出了少量驱动的查询执行框架,其中调度成为 NUMA 感知的细粒度运行时任务。碎片驱动的查询处理获取输入数据的小片段(碎片),并将这些碎片安排到运行整个运算符管道的工作线程,直到下一个管道中断。并行度没有纳入计划,但可以在查询执行期间弹性更改,因此调度程序可以对不同部分的执行速度做出反应,还可以动态调整资源以响应工作负载中新到达的查询。此外,调度程序了解 NUMA 本地部分的数据局部性和操作符状态,使得绝大多数执行发生在 NUMA 本地内存上。我们对 TPC-H 和 SSB 基准测试的评估显示出极高的绝对性能,并且在 32 个内核的情况下平均加速超过 30。
With modern computer architecture evolving, two problems conspire against the state-of-the-art approaches in parallel query execution: (i) to take advantage of many-cores, all query work must be distributed evenly among (soon) hundreds of threads in order to achieve good speedup, yet (ii) dividing the work evenly is difficult even with accurate data statistics due to the complexity of modern out-of-order cores. As a result, the existing approaches for plan-driven parallelism run into load balancing and context-switching bottlenecks, and therefore no longer scale. A third problem faced by many-core architectures is the decentralization of memory controllers, which leads to Non-Uniform Memory Access (NUMA). In response, we present the morsel-driven query execution framework, where scheduling becomes a fine-grained run-time task that is NUMA-aware. Morsel-driven query processing takes small fragments of input data (morsels) and schedules these to worker threads that run entire operator pipelines until the next pipeline breaker. The degree of parallelism is not baked into the plan but can elastically change during query execution, so the dispatcher can react to execution speed of different morsels but also adjust resources dynamically in response to newly arriving queries in the workload. Further, the dispatcher is aware of data locality of the NUMA-local morsels and operator state, such that the great majority of executions takes place on NUMA-local memory. Our evaluation on the TPC-H and SSB benchmarks shows extremely high absolute performance and an average speedup of over 30 with 32 cores.