Accurate and complexity-effective spatial pattern prediction

Accurate and complexity-effective spatial pattern prediction
复制标题

准确且复杂有效的空间模式预测

DOI:
--
复制
发表时间:
2004
期刊:
International Symposium on High-Performance Computer Architecture
影响因子:
--
通讯作者:
Andreas Moshovos
Andreas Moshovos
中科院分区:
--
文献类型:
--
作者:
Chi F. Chen;Se;B. Falsafi;Andreas Moshovos

文献摘要

被引文献

相似文献

最近的研究表明,在程序内部和跨程序中,缓存的空间使用情况有很大的变化。不幸的是,常规缓存通常采用固定的缓存线大小来平衡空间和时间位置的开发,并避免填充过宽的带宽需求。传统缓存无法利用空间变化的导致无能导致次优性能和不必要的缓存功率耗散。我们描述了空间模式预测器(SPP),这是一种经济高效的硬件机制,可准确预测运行时空间组内的参考模式(即记忆中数据的连续区域)。可以实现准确但低成本的SPP设计的关键观察是,空间模式与高速缓存线中的指令地址和数据参考偏移良好相关。我们只需要少量的预测器内存来存储预测模式。具有64字节线的64 kbyte 2-way 2-way Set-sed-sassociative LL数据缓存的仿真结果表明:(1)256个无标签的直接映射SPP平均可以实现95%的预测覆盖率。 ,过度预测的模式仅减少了8%,(2)假设使用了70 nm的过程技术,SPP平均有助于将基本缓存中的泄漏能量降低41%,造成少于1%的绩效降解,并且(3)预取用使用SPP高达512个字节的空间组平均将执行时间提高了33%,最高为两个。
Recent research suggests that there are large variations in a cache's spatial usage, both within and across programs. Unfortunately, conventional caches typically employ fixed cache line sizes to balance the exploitation of spatial and temporal locality, and to avoid prohibitive cache fill bandwidth demands. The resulting inability of conventional caches to exploit spatial variations leads to suboptimal performance and unnecessary cache power dissipation. We describe the spatial pattern predictor (SPP), a cost-effective hardware mechanism that accurately predicts reference patterns within a spatial group (i.e., a contiguous region of data in memory) at runtime. The key observation enabling an accurate, yet low-cost, SPP design is that spatial patterns correlate well with instruction addresses and data reference offsets within a cache line. We require only a small amount of predictor memory to store the predicted patterns. Simulation results for a 64-Kbyte 2-way set-associative Ll data cache with 64-byte lines show that: (1) a 256-entry tag-less direct-mapped SPP can achieve, on average, a prediction coverage of 95%, over-predicting the patterns by only 8%, (2) assuming a 70 nm process technology, the SPP helps reduce leakage energy in the base cache by 41% on average, incurring less than 1% performance degradation, and (3) prefetching spatial groups of up to 512 bytes using SPP improves execution time by 33% on average and up to a factor of two.