Register Flush-free Runahead Execution for Modern Vector Processors

Register Flush-free Runahead Execution for Modern Vector Processors
复制标题

现代矢量处理器的寄存器免刷新超前执行

DOI:
10.1109/sbac-pad53543.2021.00023
复制
发表时间:
2021
期刊:
Proceedings of 2021 IEEE 33rd International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD)
影响因子:
--
通讯作者:
Kobayashi Hiroaki
Kobayashi Hiroaki
中科院分区:
--
文献类型:
--
作者:
Takayashiki Hikaru;Sato Masayuki;Komatsu Kazuhiko;Kobayashi Hiroaki

文献摘要

相似文献

现代向量处理器被设计为实现高持续性能,特别是在HPC应用中,因为它们面向数据级并行的强大指令集。此外,最新的向量处理器采用向量指令的乱序执行,以利用由于向量算术指令和向量加载/存储指令之间的延迟的显著差距而导致的并行级并行性。尽管努力,这个差距仍然带来了现代向量处理器的持续性能的恶化。本文提出了一种用于现代向量处理器的runahead执行机制,通过进一步开发并行级并行来填补延迟间隙。如果处理器由于长等待时间指令而停顿,则传统的超前运行执行机制将处理器状态从正常模式改变为超前运行模式,并且处理器推测性地执行可能导致停顿及其依赖性的后续指令。然而,传统的超前运行执行机制在完成该模式之后刷新在超前运行模式中计算的寄存器的值,并且不能在随后的正常模式中重新使用它们。由于向量处理器即使在一个向量寄存器中也具有许多值,因此这些刷新和重新执行浪费了核心和高速缓存之间的带宽。因此,为了解决传统的Runahead机制的这个问题,我们提出的机制将包含结果的寄存器保留在Runahead模式中,以便处理器即使在返回到正常模式之后也可以使用寄存器。为了在退出runahead模式后正确使用这些寄存器,所提出的机制新实现了将runahead执行指令的提交顺序信息和寄存器别名信息继承到正常模式的功能。评估结果表明,所提出的机制提高了20%和3%的平均性能由传统的机制。
Modern vector processors have been designed to achieve high sustained performance, especially in HPC applications, because of their powerful instruction set oriented to data-level parallelism. Additionally, the latest vector processor adopts the out-of-order execution of the vector instructions to exploit instruction-level parallelism due to a significant gap in latency between vector arithmetic instructions and vector load/store instructions. In spite of the effort, this gap still brings a deterioration of sustained performance of the modern vector processors. This paper proposes a runahead execution mechanism for the modern vector processors to fill the latency gap by further exploiting instruction-level parallelism. If the processor stalls due to a long latency instruction, the conventional runahead execution mechanism changes the processor state from a normal mode to a runahead mode, and the processor speculatively executes the subsequent instructions that can cause stalls and their dependencies. However, the conventional runahead execution mechanisms flush the registers' values calculated in the runahead mode after finishing this mode and cannot reuse them in the subsequent normal mode. Since the vector processors have many values even in one vector register, these flushes and re-executions waste the bandwidth between cores and caches. Thus, to solve this problem of the conventional runahead mechanism, our proposed mechanism leaves the registers containing the results in the runahead mode in order for the processor to use the registers even after returning to the normal mode. For correctly using these registers after exiting the runahead mode, the proposed mechanism newly realizes functions to inherit the commit order information and the register aliasing information of the runahead-executed instructions into the normal mode. The evaluation results show that the proposed mechanism improves the performance by up to 20% and 3% on average by the conventional mechanism.