A 5GHz 7nm L1 cache memory compiler for high-speed computing and mobile applications

A 5GHz 7nm L1 cache memory compiler for high-speed computing and mobile applications
复制标题

适用于高速计算和移动应用的 5GHz 7nm L1 高速缓存编译器

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Solid-State Circuits Conference
影响因子:
--
通讯作者:
Jonathan Chang
Jonathan Chang
中科院分区:
--
文献类型:
--
作者:
M. Clinton;R. Singh;Marty Tsai;Shayan Zhang;B. Sheffield;Jonathan Chang

文献摘要

被引文献

相似文献

在高性能计算(HPC)应用中,L1高速缓存的速度通常决定处理器内核的最大频率(/Max)。大规模生产高性能微处理器的公司通常会使用完全自定义的宏组成L1缓存:以确保L1缓存的性能不会限制处理器的fMAX或吞吐量。此外,自定义L1缓存设计通常使用双端口8 T或大型6 T位单元,沿着多米诺读取逻辑和非常短的BL [2,3]。这些设计在密度和面积之间进行权衡以获得高性能。本文提出了一种不同的方法,它可以满足一系列不同的应用程序,内存编译器,可以产生超过10,000个不同的高速L1缓存宏配置。本文描述的7 nm L1缓存编译器使用高电流(HC)6 T位单元,其面积效率高于8 T位单元。沿着小信号感测的HC位单元允许长BL(256 b),从而导致进一步的面积效率改进。由于这些L1宏在移动的应用中的使用可能性与在HPC应用中的使用一样大,因此它们使用阵列双轨(ADR)架构实现[4]。ADR架构(图11.3.1)允许L1宏的外围电路在与处理器内核相同的电压下工作:较低的I/DD导致动态功耗节省。当SRAM和逻辑电源等效时,ADR性能也会在接口双轨上得到改善,因为ADR设计不会在输入或输出上受到电平移位器延迟的影响。
In high performance computing (HPC) applications, the speed of the L1 cache will typically determine the maximum frequency (/Max) of the processor core. Companies that mass produce high-performance microprocessors commonly have the L1 cache consist of fully-custom macros: to ensure that the performance of the L1 cache does not limit the fMAX or throughput of the processor. In addition, it is also common for the custom L1 cache designs to use a two-port 8T or a large 6T bitcell, along with domino read logic and very short BL [2,3]. These designs tradeoff density and area for high performance. This paper presents a different approach, one which can satisfy a range of different applications; a memory compiler that can generate more than 10,000 different high-speed L1 cache macro configurations is proposed. The 7nm L1-cache compiler described in this paper uses a high-current (HC) 6T bitcell, which is more area efficient than an 8T bitcell. The HC bitcell, along with small-signal sensing, allows for long BL (256b), leading to further area efficiency improvements. Since these L1 macros are just as likely to be used in mobile applications as they are to be used in HPC applications, they were implemented using the array dual-rail (ADR) architecture [4]. The ADR architecture (Fig. 11.3.1) allows the periphery circuits of the L1 macro to operate at the same voltage as the processor core: a lower l/DD results in dynamic power savings. ADR performance is also improved, over an interface dual-rail, when the SRAM and logic supplies are equivalent, as ADR design does not suffer from a level-shifter delays on the inputs or outputs.