A 4R2W register file for a 2.3GHz wire-speed POWER™ processor with double-pumped write operation

A 4R2W register file for a 2.3GHz wire-speed POWER™ processor with double-pumped write operation
复制标题

DOI:
10.1109/isscc.2011.5746308
复制
发表时间:
2011-04
期刊:
2011 IEEE International Solid-State Circuits Conference
影响因子:
--
通讯作者:
G. Ditlow;R. Montoye;S. Storino;S. M. Dance;S. Ehrenreich;Bruce M. Fleischer;T. Fox;Kyle M. Holmes;Junichi Mihara;Yutaka Nakamura;S. Onishi;Robert Shearer;D. Wendel;Leland Chang
G. Ditlow;R. Montoye;S. Storino;S. M. Dance;S. Ehrenreich;Bruce M. Fleischer;T. Fox;Kyle M. Holmes;Junichi Mihara;Yutaka Nakamura;S. Onishi;Robert Shearer;D. Wendel;Leland Chang
中科院分区:
其他
文献类型:
--
作者:
G. Ditlow;R. Montoye;S. Storino;S. M. Dance;S. Ehrenreich;Bruce M. Fleischer;T. Fox;Kyle M. Holmes;Junichi Mihara;Yutaka Nakamura;S. Onishi;Robert Shearer;D. Wendel;Leland Chang

文献摘要

被引文献

相似文献

在多端口寄存器文件中,由于字行和位行连接,内存单元大小随端口总数呈二次增长。因此,减少内存单元中的物理访问端口数量可以显著节省面积和功耗,并改善延迟。双泵浦寄存器文件在一个时钟周期内操作访问端口两次,通过将存储单元中的物理端口数量减半来减少面积——这种技术通常局限于低频应用。在单独的阵列中复制一个内存单元可以使每次拷贝中的物理读端口数量减半。在这项工作中,双泵浦写端口和复制读端口应用于高性能微处理器产品[1]的4R2W寄存器文件。本文详细介绍了该阵列的具体实现和实测硬件特性,并给出了一种快速纠错方案。所使用的技术平衡了高效率和低延迟,因此与以前的工作不同,在以前的工作中,双泵浦端口执行写入,然后读取非常大的寄存器文件[2],或者双泵浦端口没有单元级读端口减少[3]。
In multi-ported register files, memory cell size grows quadratically with the total number of ports due to wordline and bitline wiring. Reducing the number of physical access ports in a memory cell can thus lead to significant area and power savings as well as latency improvement. Double-pumped register files operate access ports twice in a single clock period to reduce area by halving the number of physical ports in the memory cell — a technique often confined to low-frequency applications. Replication of a memory cell in separate arrays halves the number of physical read ports in each copy. In this work, double-pumped write ports and replicated read ports are applied to a 4R2W register file in a highperformance microprocessor product [1]. This paper describes detailed implementation and measured hardware characteristics of this array and demonstrates a fast error correction scheme. The techniques used balance high efficiency and low latency and thus differ from previous work, in which double-pumped ports perform a write followed by a read in a very large register file [2] or where write ports are double-pumped without cell-level read port reduction [3].