2.3 A 220GOPS 96-Core Processor with 6 Chiplets 3D-Stacked on an Active Interposer Offering 0.6ns/mm Latency, 3Tb/s/mm2 Inter-Chiplet Interconnects and 156mW/mm2@ 82%-Peak-Efficiency DC-DC Converters
2.3 A 220GOPS 96-Core Processor with 6 Chiplets 3D-Stacked on an Active Interposer Offering 0.6ns/mm Latency, 3Tb/s/mm2 Inter-Chiplet Interconnects and 156mW/mm2@ 82%-Peak-Efficiency DC-DC Converters
复制标题
2.3%20A%20220GOPS%2096-Core%20处理器%20with%206%20Chiplet%203D-Stacked%20on%20an%20Active%20Interposer%20Offering%200.6ns/mm%20延迟,%203Tb/s/mm2%20Inter-Chiplet%
DOI:
10.1109/isscc19947.2020.9062927
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
F. Clermidy
中科院分区:
文献类型:
--
作者:
P. Vivet;Eric Guthmuller;Y. Thonnart;G. Pillonnet;G. Moritz;I. Miro;C. F. Tortolero;J. Durupt;C. Bernard;D. Varreau;Julian J. H. Pontes;S. Thuries;David Coriat;M. Harrand;D. Dutoit;D. Lattard;L. Arnaud;J. Charbonnier;P. Coudrain;A. Garnier;F. Berger;A. Gueugnot;A. Greiner;Quentin L. Meunier;A. Farcy;A. Arriordaz;S. Chéramy;F. Clermidy
In the context of high-performance computing and big-data applications, the quest for performance requires modular, scalable, energy-efficient, low-cost manycore systems. Partitioning the system into multiple chiplets 3D-stacked onto large-scale interposers - organic substrate [1], 2.5D passive interposer [2] or silicon bridge [3] -leads to large modular architectures and cost reductions in advanced technologies by the Known Good Die (KGD) strategy and yield management. However, these approaches lack flexible efficient long-distance communications, smooth integration of heterogeneous chiplets, and easy integration of less-scalable analog functions, such as power management [4] and system IOs. To tackle these issues, this paper presents an active interposer integrating: i) a Switched Capacitor Voltage Regulator (SCVR) for on-chip power management; ii) flexible system interconnect topologies between all chiplets for scalable cache coherency support; iii) energy-efficient 3D-plugs for dense inter-layer communication; iv) a memory-IO controller and PHY for socket communication. The chip (Fig. 2.3.7) integrates 96 cores in 6 chiplets in 28nm FDSOI CMOS, 30-stacked in a face-to-face configuration using 20µm-pitch micro-bumps (µ-bumps) onto a 200 mm2 active interposer with 40µm-pitch Through Silicon Via (TSV) middle in a 65nm technology node. Even though complex functions are integrated, active-interposer yield is high thanks to the mature 65nm node and a reduced complexity (0.08transistors/µm2), with 30% of interposer area devoted to a SCVR variability-tolerant capacitors scheme.