Slipstream Memory Hierarchies

Slipstream Memory Hierarchies
复制标题

滑流内存层次结构

DOI:
--
复制
发表时间:
2002
期刊:
影响因子:
--
通讯作者:
E. Rotenberg
E. Rotenberg
中科院分区:
--
文献类型:
--
作者:
Zachary Purser;K. Sundaramoorthy;E. Rotenberg

文献摘要

被引文献

相似文献

Slipstream处理器在CHIP多校区(CMP)中利用了其他未使用的处理元件,以加快单个程序的速度。它通过运行该程序的两个冗余副本来做到这一点。预测的非注重计算是从其中一个程序中删除的,从而加快了它。第二个程序检查第一个程序的前进进度,并且在此过程中也加速了。这两个程序副本都比任何一个人都更早完成。冗余程序在架构上是独立的,这导致了一个简单的执行模型。物理内存页面由操作系统重复,使处理器明确管理固定量的透明命名存储。不幸的是,将内存使用量加倍会部分否定性能提高。我们观察到1)CMP中已经复制的L1缓存提供了足够的隐式重命名存储,并且2)此存储不需要明确管理,因为Slipstream Paradigm耐受了稍微不准确的内存重命名。利用典型的私有L1/共享L2内存层次结构中的未修改的缓存操作,我们开发了一种有效的基于硬件的内存重复方法,该方法极大地超过了基于软件的重复,但不需要任何明确的硬件管理。此外,新的重复方法可以在投机计划分歧时更简单的状态恢复。简单的高速缓存冲洗消除了先前所需的Slipstream恢复组件。通过利用冲洗式高速缓存线中的保留数据作为高度准确的价值预测,可以降低冲洗引起的强制性失误的性能影响。
A slipstream processor harnesses an otherwise unused processing element in a chip multipro-cessor (CMP) to speed up a single program. It does this by running two redundant copies of the program. Predicted-non-essential computation is speculatively removed from one of the programs , speeding it up. The second program checks the forward progress of the first and is also sped up in the process. Both program copies finish sooner than either can alone. The redundant programs are architecturally independent and this leads to a simple execution model. Physical memory pages are duplicated by the operating system, sparing the processor from explicitly managing a fixed amount of transparent rename storage. Unfortunately, doubling memory usage partially negates performance gains. We observe that 1) the already-replicated L1 caches in a CMP provide enough implicit rename storage and 2) this storage does not need to be explicitly managed because the slipstream paradigm is tolerant of slightly inaccurate memory renaming. Leveraging unmodified cache actions within a typical private-L1/shared-L2 memory hierarchy, we develop an efficient hardware-based memory duplication approach that significantly outperforms software-based duplication, yet does not require any explicit hardware management. Furthermore, the new duplication approach enables much simpler state recovery when the speculative program diverges. Simple cache flushing eliminates a previously-required slipstream recovery component. And the performance impact of flush-induced compulsory misses is reduced by exploiting preserved data within flushed cache lines as highly-accurate value predictions.