Memory Architecture for Integrating Emerging Memory Technologies

Memory Architecture for Integrating Emerging Memory Technologies
复制标题

用于集成新兴内存技术的内存架构

DOI:
--
复制
发表时间:
2011
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Zhichun Zhu
Zhichun Zhu
中科院分区:
--
文献类型:
--
作者:
Kun Fang;Long Chen;Zhao Zhang;Zhichun Zhu

文献摘要

被引文献

相似文献

当前的主存储系统设计受到数十年历史的同步DRAM体系结构的严重限制,该体系结构需要存储器控制器跟踪内存设备(芯片)的内部状态并安排所有设备操作的时机。这种刚度已成为将PCM等新兴记忆技术(例如PCM)整合到现有内存系统中的障碍,因为它们的时序要求大不相同。此外,随着将存储器控制器嵌入处理器中的趋势,至关重要的是在通用处理器和多种内存模块之间具有互操作性。为了解决这个问题,我们提出了一个新的内存体系结构框架,称为通用内存体系结构(UNIMA)。它通过使用每个存储器模块的桥接芯片来执行本地调度,从而可以通过将设备操作的调度分解为互操作性。新的体系结构还可以帮助提高内存可扩展性,功率效率和带宽,如先前建议的解耦内存组织。这项研究的重点是评估设备操作本地调度的性能影响。我们在DDRX内存总线上提出了UNIMA的原型实现,然后通过不同的工作负载评估其效率。仿真结果表明,由于内存模块之间的并行性增加,UNIMA实际上提高了内存密集型工作负载的内存系统效率。传统DDRX内存体系结构的总体性能改善平均为3.1%。由于记忆潜伏期的少量增加,其他工作负载的性能平均略有减少1.0%。简而言之,原型和评估表明,可以将各种内存技术整合到单个内存架构中,几乎不会损失整体性能。
Current main memory system design is severely limited by the decades-old synchronous DRAM architecture, which requires the memory controller to track the internal status of memory devices (chips) and schedule the timing of all device operations. This rigidity has become an obstacle of integrating emerging memory technologies such as PCM into existing memory systems, because their timing requirements are vastly different. Furthermore, with the trend of embedding memory controllers into processors, it is crucial to have interoperability among general-purpose processors and diverse memory modules. To address this issue, we propose a new memory architecture framework called universal memory architecture (UniMA). It enables the interoperability by decoupling the scheduling of device operations from memory controller, using a bridge chip at each memory module to perform local scheduling. The new architecture may also help improve memory scalability, power efficiency, and bandwidth as previously proposed decoupled memory organizations. A major focus of this study is to evaluate the performance impact of local scheduling of device operations. We present a prototype implementation of UniMA on top of DDRx memory bus, and then evaluate its efficiency with different workloads. The simulation results show that UniMA actually improves memory system efficiency for memory-intensive workloads due to increased parallelism among memory modules. The overall performance improvement over the conventional DDRx memory architecture is 3.1% on average. The performance of other workloads is reduced slightly, by 1.0% on average, due to a small increase of memory idle latency. In short, the prototype and evaluation demonstrate that it is possible to integrate diverse memory technologies into a single memory architecture with virtually no loss of overall performance.