Synergistic circuit and system design for energy-efficient and robust domain wall caches

Synergistic circuit and system design for energy-efficient and robust domain wall caches
复制标题

用于节能且稳健的畴壁缓存的协同电路和系统设计

DOI:
10.1145/2627369.2627643
复制
发表时间:
2014
期刊:
2014 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED)
影响因子:
--
通讯作者:
Swaroop Ghosh
Swaroop Ghosh
中科院分区:
--
文献类型:
--
作者:
Seyedhamidreza Motaman;Anirudh Iyengar;Swaroop Ghosh

文献摘要

被引文献

相似文献

非易失性存储器由于其低待机功耗和出色的保持能力而在嵌入式缓存应用中获得了极大的关注。畴壁存储器(DWM)是一种可能的候选者,这是由于其能够在每个单元中存储多个位以便打破密度屏障。此外,它还提供低待机功耗、快速访问时间、良好的耐用性和保持力。然而,其遭受差的写入等待时间、移位等待时间、移位功率和写入功率。DWM本质上是顺序的,读/写操作的延迟取决于读/写头的位偏移。本文研究了电路设计的挑战,如位单元布局,头部定位,纳米线的利用率,移位功率,移位延迟,并提供解决方案,以处理这些问题。提出了一种协同系统,该系统通过将诸如合并读/写头(用于紧凑布局)、翻转位单元和移位门控(用于移位功率优化)、WL(WL)捆绑(用于访问延迟)、移位电路设计等电路技术与诸如分段高速缓存等微架构技术相结合来实现节能且鲁棒的DWM高速缓存。仿真结果显示,在广泛的PARSEC基准测试中,性能提高了3-33%,功耗提高了1.25 - 14.4倍。
Non-volatile memories are gaining significant attention for embedded cache application due to their low standby power and excellent retention. Domain wall memory (DWM) is one possible candidate due to its ability to store multiple bits per cell in order to break the density barrier. Additionally, it provides low standby power, fast access time, good endurance and retention. However, it suffers from poor write latency, shift latency, shift power and write power. DWM is sequential in nature and latency of read/write operations depends on the offset of the bit from the read/write head. This paper investigates the circuit design challenges such as bitcell layout, head positioning, utilization factor of the nanowire, shift power, shift latency and provides solutions to deal with these issues. A synergistic system is proposed by combining circuit techniques such as merged read/write heads (for compact layout), flipped-bitcell and shift gating (for shift power optimization), wordline (WL) strapping (for access latency), shift circuit design with micro-architectural techniques such as segmented cache to realize energy-efficient and robust DWM cache. Simulations show 3-33% better performance and 1.25X-14.4X better power over a wide range of PARSEC benchmarks.