Precise Runahead Execution

Precise Runahead Execution
复制标题

精确的超前执行

DOI:
--
复制
发表时间:
2020
期刊:
International Symposium on High-Performance Computer Architecture
影响因子:
--
通讯作者:
L. Eeckhout
L. Eeckhout
中科院分区:
--
文献类型:
--
作者:
Ajeya Naithani;Josué Feliu;Almutaz Adileh;L. Eeckhout

文献摘要

参考文献

被引文献

相似文献

RunAhead执行通过准确预取长期内存访问来改善处理器性能。当长期负载导致指令窗口填充和停止管道时,处理器进入Runahead模式,并保持投机执行代码以触发准确的预取。最近的改进跟踪导致长期负载的指令链,将其存储在Runahead Buffer中,并在Runahead执行期间仅执行此链,目的是生成更多的预摘要请求。不幸的是,所有先前的Runahead提案都有限制性能和能源效率的缺点,因为它们在输入Runahead模式时会释放处理器状态,然后需要对管道进行补充以重新启动正常操作。此外,Runahead Buffer仅通过仅跟踪一条指令链,从而限制预取覆盖范围,该指令链导致相同的长期负载。我们提出了精确的RunAhead执行(PRE),该执行(PRE)基于关键观察,即输入RunaHead模式时,处理器具有足够的问题队列和物理寄存器文件资源来投机执行指令。这减轻了在ROB,发行队列和物理寄存器文件中释放和重新填充处理器状态的需求。此外,预先执行前的预先执行以下指令,以导致全窗户摊位,使用新颖的寄存器重命名机制在Runahead模式下快速释放物理寄存器,从而进一步提高效率和有效性。最后,可选的预示器在前端中解码的Runahead微型鞋,以节省能量。我们使用一组记忆密集型应用程序进行的实验评估表明,与最近的Runahead提案相比,PRE可以额外提高18.2%的性能,同时将能耗降低6.8%。
Runahead execution improves processor performance by accurately prefetching long-latency memory accesses. When a long-latency load causes the instruction window to fill up and halt the pipeline, the processor enters runahead mode and keeps speculatively executing code to trigger accurate prefetches. A recent improvement tracks the chain of instructions that leads to the long-latency load, stores it in a runahead buffer, and executes only this chain during runahead execution, with the purpose of generating more prefetch requests. Unfortunately, all prior runahead proposals have shortcomings that limit performance and energy efficiency because they release processor state when entering runahead mode and then need to refill the pipeline to restart normal operation. Moreover, runahead buffer limits prefetch coverage by tracking only a single chain of instructions that leads to the same long-latency load. We propose precise runahead execution (PRE) which builds on the key observation that when entering runahead mode, the processor has enough issue queue and physical register file resources to speculatively execute instructions. This mitigates the need to release and re-fill processor state in the ROB, issue queue, and physical register file. In addition, PRE pre-executes only those instructions in runahead mode that lead to full-window stalls, using a novel register renaming mechanism to quickly free physical registers in runahead mode, further improving efficiency and effectiveness. Finally, PRE optionally buffers decoded runahead micro-ops in the frontend to save energy. Our experimental evaluation using a set of memory-intensive applications shows that PRE achieves an additional 18.2% performance improvement over the recent runahead proposals while at the same time reducing energy consumption by 6.8%.
R3-DLA(减少、重用、回收):一种更有效的解耦前瞻架构方法
DOI: 10.1109/hpca.2019.00064
发表时间: 2019
期刊: 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA
影响因子: --
作者:
Kondguli, Sushant;Huang, Michael
通讯作者: Huang, Michael