Effective ahead pipelining of instruction block address generation

Effective ahead pipelining of instruction block address generation
复制标题

指令块地址生成的有效提前流水线

DOI:
--
复制
发表时间:
2003
期刊:
30th Annual International Symposium on Computer Architecture, 2003. Proceedings.
影响因子:
--
通讯作者:
A. Fraboulet
A. Fraboulet
中科院分区:
--
文献类型:
--
作者:
André Seznec;A. Fraboulet

文献摘要

被引文献

相似文献

在N路发布超标量处理器上,前端取指令引擎必须以高于每周期N条指令的持续速率将指令递送到执行核心。这意味着指令地址生成器/预测器(IAG)必须以更高的速率预测指令流,同时不能牺牲预测精度。实现这种预测的高精度变得越来越重要,因为随着每一代新处理器的出现,整个流水线变得越来越深。然后使用非常复杂的IAG,其具有用于跳转、返回、条件和无条件分支以及复杂逻辑的不同预测器。通常,IAG使用信息(分支历史、提取地址...)在一个周期内可用于预测下一个提取地址。不幸的是,复杂的IAG无法在短周期内提供预测。因此,处理器依赖于IAG的层次结构,其准确性不断提高,但延迟也不断增加:准确但缓慢的IAG用于纠正快速但不太准确的IAG。由于这些校正,潜在指令带宽的很大一部分通常浪费在流水线气泡中。作为使用IAG层次结构的替代方案,可以在其使用之前的几个周期启动指令地址生成。我们详细探讨了这种先进的流水线IAG。所示的示例示出,即使当指令地址生成在其使用之前(部分地)启动五个周期时,也可以达到与传统的一个块前复合IAG的预测精度大致相同的预测精度。所提出的解决方案允许提供接近于每周期一个指令块的持续地址生成速率,具有最先进的精度。
On a N-way issue superscalar processor, the front end instruction fetch engine must deliver instructions to the execution core at a sustained rate higher than N instructions per cycle. This means that the instruction address generator/predictor (IAG) has to predict the instruction flow at an even higher rate while the prediction accuracy cannot be sacrificed. Achieving high accuracy on this prediction becomes more and more critical since the overall pipeline is becoming deeper and deeper with each new generation of processors. Then very complex IAGs featuring different predictors for jumps, returns, conditional and unconditional branches and complex logic are used. Usually, the IAG uses information (branch histories, fetch addresses, ...) available at a cycle to predict the next fetch address(es). Unfortunately, a complex IAG cannot deliver a prediction within a short cycle. Therefore, processors rely on a hierarchy of IAGs with increasing accuracies but also increasing latencies: the accurate but slow IAG is used to correct the fast, but less accurate IAG. A significant part of the potential instruction bandwidth is often wasted in pipeline bubbles due to these corrections. As an alternative to the use of a hierarchy of IAGs, it is possible to initiate the instruction address generation several cycles ahead of its use. We explore in details such an ahead pipelined IAG. The example illustrated shows that, even when the instruction address generation is (partially) initiated five cycles ahead of its use, it is possible to reach approximately the same prediction accuracy as the one of a conventional one block ahead complex IAG. The solution presented allows to deliver a sustained address generation rate close to one instruction block per cycle with state of the art accuracy.