2.3 A 220GOPS 96-Core Processor with 6 Chiplets 3D-Stacked on an Active Interposer Offering 0.6ns/mm Latency, 3Tb/s/mm2 Inter-Chiplet Interconnects and 156mW/mm2@ 82%-Peak-Efficiency DC-DC Converters

2.3 A 220GOPS 96-Core Processor with 6 Chiplets 3D-Stacked on an Active Interposer Offering 0.6ns/mm Latency, 3Tb/s/mm2 Inter-Chiplet Interconnects and 156mW/mm2@ 82%-Peak-Efficiency DC-DC Converters
复制标题

2.3%20A%20220GOPS%2096-Core%20处理器%20with%206%20Chiplet%203D-Stacked%20on%20an%20Active%20Interposer%20Offering%200.6ns/mm%20延迟,%203Tb/s/mm2%20Inter-Chiplet%

DOI:
10.1109/isscc19947.2020.9062927
复制
发表时间:
2020
期刊:
2020 IEEE International Solid- State Circuits Conference - (ISSCC)
影响因子:
--
通讯作者:
F. Clermidy
F. Clermidy
中科院分区:
--
文献类型:
--
作者:
P. Vivet;Eric Guthmuller;Y. Thonnart;G. Pillonnet;G. Moritz;I. Miro;C. F. Tortolero;J. Durupt;C. Bernard;D. Varreau;Julian J. H. Pontes;S. Thuries;David Coriat;M. Harrand;D. Dutoit;D. Lattard;L. Arnaud;J. Charbonnier;P. Coudrain;A. Garnier;F. Berger;A. Gueugnot;A. Greiner;Quentin L. Meunier;A. Farcy;A. Arriordaz;S. Chéramy;F. Clermidy

文献摘要

被引文献

相似文献

在高性能计算和大数据应用的背景下,对性能的追求需要模块化、可扩展、节能、低成本的众核系统。将系统划分为3D堆叠到大规模中介层上的多个小芯片-有机基板[1],2.5D无源中介层[2]或硅桥[3] -通过已知良好管芯(KGD)策略和良率管理导致大型模块化架构和先进技术的成本降低。然而,这些方法缺乏灵活有效的长距离通信,异构小芯片的平滑集成,以及可扩展性较低的模拟功能(如电源管理[4]和系统IO)的轻松集成。为了解决这些问题,本文提出了一种有源插入器,其集成:i)用于片上电源管理的开关电容电压调节器(SCVR); ii)用于可扩展高速缓存一致性支持的所有小芯片之间的灵活系统互连拓扑; iii)用于密集层间通信的节能3D插头; iv)用于插座通信的存储器IO控制器和PHY。该芯片(图2.3.7)在28 nm FDSOI CMOS中将96个内核集成在6个小芯片中,采用面对面配置,在65 nm技术节点中,使用20µ m间距微凸块(µ-bump)将30个内核堆叠到200 mm 2有源中介层上,中间为40µ m间距硅通孔(TSV)。尽管集成了复杂的功能,但由于成熟的65 nm节点和降低的复杂性(0.08个晶体管/µm2),30%的中介层面积专用于SCVR可变容差电容器方案,有源中介层产量很高。
In the context of high-performance computing and big-data applications, the quest for performance requires modular, scalable, energy-efficient, low-cost manycore systems. Partitioning the system into multiple chiplets 3D-stacked onto large-scale interposers - organic substrate [1], 2.5D passive interposer [2] or silicon bridge [3] -leads to large modular architectures and cost reductions in advanced technologies by the Known Good Die (KGD) strategy and yield management. However, these approaches lack flexible efficient long-distance communications, smooth integration of heterogeneous chiplets, and easy integration of less-scalable analog functions, such as power management [4] and system IOs. To tackle these issues, this paper presents an active interposer integrating: i) a Switched Capacitor Voltage Regulator (SCVR) for on-chip power management; ii) flexible system interconnect topologies between all chiplets for scalable cache coherency support; iii) energy-efficient 3D-plugs for dense inter-layer communication; iv) a memory-IO controller and PHY for socket communication. The chip (Fig. 2.3.7) integrates 96 cores in 6 chiplets in 28nm FDSOI CMOS, 30-stacked in a face-to-face configuration using 20µm-pitch micro-bumps (µ-bumps) onto a 200 mm2 active interposer with 40µm-pitch Through Silicon Via (TSV) middle in a 65nm technology node. Even though complex functions are integrated, active-interposer yield is high thanks to the mature 65nm node and a reduced complexity (0.08transistors/µm2), with 30% of interposer area devoted to a SCVR variability-tolerant capacitors scheme.