Segmenting Age Matrices to Improve Instruction Scheduling without Increasing Delay and Area

Segmenting Age Matrices to Improve Instruction Scheduling without Increasing Delay and Area
复制标题

DOI:
10.1109/iccd56317.2022.00059
复制
发表时间:
2022-10
期刊:
2022 IEEE 40th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
H. Ando
H. Ando
中科院分区:
其他
文献类型:
--
作者:
H. Ando

文献摘要

相似文献

当前超标量处理器在发布队列(IQ)中有一个称为年龄矩阵(AM)的特殊电路,它选择队列中最早的就绪指令,以允许更好的指令调度。然而,优化级别是不够的,因为AM只选择单个最旧的指令,而要发出的其他指令是随机选择的。在本文中,我们提出了一个新的AM组织和一个计划,成功地使用它,我们称之为分段AM(SegAM)。在SegAM中,AM被物理分段,因此,每个AM段比原始AM小二次方。因此,允许选择多个最旧指令的多个AM可以被插入到IQ中,其总面积和延迟保持不变或从单个单片AM的总面积和延迟减少。为了确保AM段选择整个IQ中最旧的就绪指令,该方案分派(即,写入)指令到IQ逐段,其按年龄对段进行排序。我们使用SPEC2017基准程序的评估结果表明,具有三个AM的IQ,其中每个AM被分割为四个,比传统的单个单片AM平均提高了6.4%和1.2%(高达17.7%和11.4%)分别为整数和浮点程序,减少了21%的IQ延迟,8%的IQ面积和84%的年龄矩阵能量。
Current superscalar processors have a special circuit called the age matrix (AM) in the issue queue (IQ), which selects the oldest ready instruction in the queue, to allow better instruction scheduling. However, the optimization level is insufficient, because the AM selects only the single oldest instruction, and the other instructions to be issued are selected randomly. In this paper, we propose a new AM organization and a scheme that uses it successfully, which we call the segmented AM (SegAM). In SegAM, the AM is physically segmented, and therefore, each AM segment is quadratically smaller than the original AM. Consequently, multiple AMs, which allow the multiple oldest instructions to be selected, can be inserted into the IQ with their total area and delay remaining unchanged or reduced from that of the single monolithic AM. To ensure that an AM segment selects the oldest ready instruction in the entire IQ, the scheme dispatches (i.e., writes) instructions to the IQ segment-by-segment, which orders the segments by age. Our evaluation results using SPEC2017 benchmark programs demonstrates that an IQ with three AMs, in which each AM is segmented into four, achieves higher performance than the conventional single monolithic AM by an average of 6.4% and 1.2% (up to 17.7% and 11.4%) for integer and floating-point programs, respectively, with reductions of 21% IQ delay, 8% IQ area, and 84% age matrix energy.