Single Instruction Multiple Data – Not Everything is a Nail for this Hammer

Single Instruction Multiple Data – Not Everything is a Nail for this Hammer
复制标题

单指令多数据——并非一切都是这把锤子的钉子

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
David Broneske
David Broneske
中科院分区:
--
文献类型:
--
作者:
David Broneske

文献摘要

被引文献

相似文献

几十年来,硬件供应商一直在努力与电源和内存墙作斗争[1,2]。由于大部分处理时间取决于指令的数量、所使用的寄存器的数量以及指令之间的依赖性,而不取决于寄存器的大小,因此向量的独立数据项(即,一列)可以并行处理。因此,一个银线似乎是单指令多数据(SIMD)-一种在当前CPU上可用的处理范例,但也有加速卡,如GPU和MIC(即,Intel Xeon Phi)。例如,聚合可以并行地对多个数据项执行求和或计数指令。通过在128位SSE寄存器中加载四个32位整数,并在一个周期内对所有四个数据项执行加法,应该可以获得四倍的性能优势。然而,这些高期望在实践中很少得到满足。在这次演讲中,我们详细介绍了我们在使用SIMD优化数据库操作符时遇到的陷阱。总的来说,这些陷阱可以在不同的级别上找到:特别是操作员内部的数据移动和数据布局对性能改进起着至关重要的作用。
Hardware vendors have been struggling to fight the power and memory wall for decades [1, 2]. Since most of the processing time depends on the number of instructions, number of used registers and dependencies between instructions, but not on the size of a register, independent data items of a vector (i.e., a column) could be processed in parallel. Hence, a silver lining seems to be Single Instruction Multiple Data (SIMD) – a processing paradigm available on current CPUs, but also accelerator cards such as GPUs and MICs (i.e., Intel Xeon Phi). For instance, aggregations could perform sum or count instructions on several data items in parallel. By loading four 32-bit integers in an 128-bit SSE register and performing the addition in one cycle for all four data items, a four-fold performance benefit should be possible. However, these high expectations are rarely met in practice. In this talk, we elaborate about pitfalls that we encountered while optimizing database operators with SIMD. Overall, these pitfalls can be found at different levels: especially the data movement within the operator and the data layout plays a vital role for the performance improvements.