Optimizing main-memory join on modern hardware

Optimizing main-memory join on modern hardware
复制标题

DOI:
10.1109/tkde.2002.1019210
复制
发表时间:
2002-07-01
影响因子:
8.9
通讯作者:
Kersten, M
Kersten, M
中科院分区:
计算机科学2区
文献类型:
--
作者:
Manegold, S;Boncz, P;Kersten, M

文献摘要

被引文献

相似文献

在过去的十年中,商品CPU速度的指数级增长远远超过了内存延迟的增长。第二个趋势是,CPU性能的提高不仅是通过提高时钟速率来实现的,而且还通过提高CPU内部的并行性来实现。当前的数据库系统还没有适应这些趋势,并且在当前硬件上显示出CPU和存储器资源的不良利用率。在本文中,我们将展示如何针对大型连接优化这些资源,并将这些见解转化为未来数据库架构的指南,包括数据结构,算法,成本建模和实现。特别是,我们将讨论如何垂直碎片化的数据结构优化缓存性能的顺序数据访问。在算法方面,我们改进了分区的哈希连接与一个新的分区算法称为基数集群,这是专门设计来优化内存访问。该算法的性能量化使用一个详细的分析模型,结合内存访问成本方面的有限数量的参数,如高速缓存的大小和未命中的处罚。我们还提出了一个校准。从任何计算机硬件中自动提取此类参数的工具。我们的模型的准确性证明了详尽的实验与莫奈数据库系统在三个不同的硬件平台。最后,我们调查的实现技术,优化CPU资源的使用效果。我们的实验表明,大连接可以加速几乎一个数量级的现代RISC硬件时,内存和CPU资源都得到优化。
In the past decade, the exponential growth in commodity CPU's speed has far outpaced advances in memory latency. A second trend is that CPU performance advances are not only brought by increased clock rate, but also by increasing parallelism inside the CPU. Current database systems have not yet adapted to these trends and show poor utilization of both CPU and memory resources on current hardware. In this paper, we show how these resources can be optimized for large joins and translate these insights into guidelines for future database architectures, encompassing data structures, algorithms, cost modeling, and implementation. In particular, we discuss how vertically fragmented data structures optimize cache performance on sequential data access. On the algorithmic side, we refine the partitioned hash-join with a new partitioning algorithm called radix-cluster, which is specifically designed to optimize memory access. The performance of this algorithm is quantified using a detailed analytical model that incorporates memory access costs in terms of a limited number of parameters, such as cache sizes and miss penalties. We also present a calibration. tool that extracts such parameters automatically from any computer hardware. The accuracy of our models is proven by exhaustive experiments conducted with the Monet database system on three different hardware platforms. Finally, we investigate the effect of implementation techniques that optimize CPU resource usage. Our experiments show that large joins can be accelerated almost an-order of magnitude on modern RISC hardware when both memory and CPU resources are optimized.