Compiler-assisted preferred caching for embedded systems with STT-RAM based hybrid cache

Compiler-assisted preferred caching for embedded systems with STT-RAM based hybrid cache
复制标题

DOI:
10.1145/2248418.2248434
复制
发表时间:
2012-05
期刊:
--
影响因子:
--
通讯作者:
Qing'an Li;Mengying Zhao;C. Xue;Yanxiang He
Qing'an Li;Mengying Zhao;C. Xue;Yanxiang He
中科院分区:
其他
文献类型:
--
作者:
Qing'an Li;Mengying Zhao;C. Xue;Yanxiang He

文献摘要

被引文献

相似文献

随着技术规模的缩小,能耗正在成为传统的基于SRAM的缓存层次结构的一个大问题。新兴的自旋扭矩转移RAM(STT-RAM)由于其超低的漏电流功耗和高存储密度而成为大容量片上缓存的理想替代品。然而,STT-RAM上的写操作遭受比SRAM高得多的能量消耗和更长的延迟。混合高速缓存组成的SRAM和STT-RAM最近已提出的性能和能源效率。大多数混合缓存的管理策略采用基于迁移的技术,动态地将写密集型数据从STT-RAM移动到SRAM。这些技术导致额外的开销。在本文中,我们提出了一种编译器辅助的方法,首选缓存,以显着减少迁移开销,迁移密集型内存块的首选SRAM部分的混合高速缓存。此外,提出了一种数据分配技术,以提高优先缓存的效率。迁移开销的减少可以反过来提高基于STT-RAM的混合缓存的性能和能量效率。实验结果表明,与所提出的技术,平均而言,迁移的数量减少了21.3%,总延迟减少了8.0%,总的动态能量减少了10.8%。
As technology scales down, energy consumption is becoming a big problem for traditional SRAM-based cache hierarchies. The emerging Spin-Torque Transfer RAM (STT-RAM) is a promising replacement for large on-chip cache due to its ultra low leakage power and high storage density. However, write operations on STT-RAM suffer from considerably higher energy consumption and longer latency than SRAM. Hybrid cache consisting of both SRAM and STT-RAM has been proposed recently for both performance and energy efficiency. Most management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques lead to extra overheads. In this paper, we propose a compiler-assisted approach, preferred caching, to significantly reduce the migration overhead by giving migration-intensive memory blocks the preference for the SRAM part of the hybrid cache. Furthermore, a data assignment technique is proposed to improve the efficiency of preferred caching. The reduction of migration overhead can in turn improve the performance and energy efficiency of STT-RAM based hybrid cache. The experimental results show that, with the proposed techniques, on average, the number of migrations is reduced by 21.3%, the total latency is reduced by 8.0% and the total dynamic energy is reduced by 10.8%.