Optimization of instruction fetch mechanisms for high issue rates

Optimization of instruction fetch mechanisms for high issue rates
复制标题

优化指令获取机制以实现高发出率

DOI:
10.1145/223982.224444
复制
发表时间:
1995
期刊:
Proceedings 22nd Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
B. Patel
B. Patel
中科院分区:
--
文献类型:
--
作者:
T. Conte;Kishore N. Menezes;P. Mills;B. Patel

文献摘要

被引文献

相似文献

最近的超标量处理器每个周期发出四条指令。这些处理器还配备了高度并行的超标量内核。只有当馈送高指令带宽时,才能挖掘潜在的性能。这项任务由指令提取单元负责。准确的分支预测和较低的I-高速缓存未命中率对于提取单元的有效操作至关重要。一些关于高速缓存设计和分支预测的研究解决了这个问题。然而,这些技术是不够的。即使在存在有效的高速缓存设计和分支预测的情况下,提取单元也必须连续地从指令高速缓存中提取多个非顺序指令,以正确的顺序重新排列这些指令,并将它们提供给解码器。本文探讨了这一问题的解决方案,并提出了几种不同程度的性能和成本方案。最通用的方案是折叠缓冲区,它实现了近乎完美的性能,并且在广泛的发布率范围内,在超过90%的时间内一致地对齐指令。此外,还研究了编译器优化技术带来的性能提升。结果表明,编译器优化可以显著提高所有方案的性能。由编译器技术补充的折叠缓冲区仍然是性能最好的机制。文章最后提出了一些建议和对未来的建议。
Recent superscalar processors issue four instructions per cycle. These processors are also powered by highly-parallel superscalar cores. The potential performance can only be exploited when fed by high instruction bandwidth. This task is the responsibility of the instruction fetch unit. Accurate branch prediction and low I-cache miss ratios are essential for the efficient operation of the fetch unit. Several studies on cache design and branch prediction address this problem. However, these techniques are not sufficient. Even in the presence of efficient cache designs and branch prediction, the fetch unit must continuously extract multiple, non-sequential instructions from the instruction cache, realign these in the proper order, and supply them to the decoder. This paper explores solutions to this problem and presents several schemes with varying degrees of performance and cost. The most-general scheme, the collapsing buffer, achieves near-perfect performance and consistently aligns instructions in excess of 90% of the time, over a wide range of issue rates. The performance boost provided by compiler optimization techniques is also investigated. Results show that compiler optimization can significantly enhance performance across all schemes. The collapsing buffer supplemented by compiler techniques remains the best-performing mechanism. The paper closes with recommendations and suggestions for future.