Single Instruction Multiple Data – Not Everything is a Nail for this Hammer
Single Instruction Multiple Data – Not Everything is a Nail for this Hammer
复制标题
单指令多数据——并非一切都是这把锤子的钉子
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
David Broneske
中科院分区:
文献类型:
--
作者:
David Broneske
Hardware vendors have been struggling to fight the power and memory wall for decades [1, 2]. Since most of the processing time depends on the number of instructions, number of used registers and dependencies between instructions, but not on the size of a register, independent data items of a vector (i.e., a column) could be processed in parallel. Hence, a silver lining seems to be Single Instruction Multiple Data (SIMD) – a processing paradigm available on current CPUs, but also accelerator cards such as GPUs and MICs (i.e., Intel Xeon Phi). For instance, aggregations could perform sum or count instructions on several data items in parallel. By loading four 32-bit integers in an 128-bit SSE register and performing the addition in one cycle for all four data items, a four-fold performance benefit should be possible. However, these high expectations are rarely met in practice. In this talk, we elaborate about pitfalls that we encountered while optimizing database operators with SIMD. Overall, these pitfalls can be found at different levels: especially the data movement within the operator and the data layout plays a vital role for the performance improvements.