FXA: Executing Instructions in Front-End for Energy Efficiency

FXA: Executing Instructions in Front-End for Energy Efficiency
复制标题

DOI:
10.1587/transinf.2015edp7316
复制
发表时间:
2016-04
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Ryota Shioya;R. Takami;M. Goshima;H. Ando
Ryota Shioya;R. Takami;M. Goshima;H. Ando
中科院分区:
其他
文献类型:
--
作者:
Ryota Shioya;R. Takami;M. Goshima;H. Ando

文献摘要

相似文献

乱序超标量处理器具有高性能,但动态指令调度消耗大量能量。我们提出了一种前端执行架构(FXA),用于提高无序超标量处理器的能效。 FXA 有两个执行单元:乱序执行单元 (OXU) 和中序执行单元 (IXU)。 OXU是常见的乱序超标量处理器的执行核心。相比之下,IXU 仅由功能单元和旁路网络组成。 IXU位于处理器前端,按顺序执行指令。 IXU 充当 OXU 的滤波器。获取的指令首先被送入 IXU,如果指令准备好执行,则按顺序执行。在IXU中执行的指令被从指令流水线中移除并且不在OXU中执行。 IXU不包含动态调度逻辑,因此其能耗较低。评估结果表明,FXA可以使用IXU执行超过50%的指令,从而可以在不导致性能下降的情况下缩小耗能的OXU。因此,FXA 实现了高性能和低能耗。我们评估了 FXA,并将其与 ARM big.LITTLE 架构后的传统乱序/有序超标量处理器进行了比较。结果表明,FXA 在 SPECCPU INT 2006 基准测试套件中相对于传统超标量处理器(大)实现了几何平均值 7.4% 的性能提升,同时整个处理器的能耗降低了 17%。 FXA 的性能/能量比(能量延迟乘积的倒数)比传统超标量处理器(big)高 25%,比传统有序超标量处理器(LITTLE)高 27%。关键词: 超标量处理器, 混合有序/乱序核心, 能效
Out-of-order superscalar processors have high performance but consume a large amount of energy for dynamic instruction scheduling. We propose a front-end execution architecture (FXA) for improving the energy efficiency of out-of-order superscalar processors. FXA has two execution units: an out-of-order execution unit (OXU) and an inorder execution unit (IXU). The OXU is the execution core of a common out-of-order superscalar processor. In contrast, the IXU consists only of functional units and a bypass network only. The IXU is placed at the processor front end and executes instructions in order. The IXU functions as a filter for the OXU. Fetched instructions are first fed to the IXU, and the instructions are executed in order if they are ready to execute. The instructions executed in the IXU are removed from the instruction pipeline and are not executed in the OXU. The IXU does not include dynamic scheduling logic, and thus its energy consumption is low. Evaluation results show that FXA can execute more than 50% of the instructions by using IXU, thereby making it possible to shrink the energy-consuming OXU without incurring performance degradation. As a result, FXA achieves both high performance and low energy consumption. We evaluated FXA and compared it with conventional out-of-order/in-order superscalar processors after ARM big.LITTLE architecture. The results show that FXA achieves performance improvements of 7.4% on geometric mean in SPECCPU INT 2006 benchmark suite relative to a conventional superscalar processor (big), while reducing the energy consumption by 17% in the entire processor. The performance/energy ratio (the inverse of the energy-delay product) of FXA is 25% higher than that of a conventional superscalar processor (big) and 27% higher than that of a conventional in-order superscalar processor (LITTLE). key words: superscalar processor, hybrid in-order/out-of-order core, energy efficiency