Dark silicon and the end of multicore scaling

Dark silicon and the end of multicore scaling
复制标题

DOI:
10.1145/2000064.2000108
复制
发表时间:
2011-06
期刊:
2011 38th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
H. Esmaeilzadeh;Emily R. Blem;Renée St. Amant;Karthikeyan Sankaralingam;D. Burger
H. Esmaeilzadeh;Emily R. Blem;Renée St. Amant;Karthikeyan Sankaralingam;D. Burger
中科院分区:
其他
文献类型:
--
作者:
H. Esmaeilzadeh;Emily R. Blem;Renée St. Amant;Karthikeyan Sankaralingam;D. Burger

文献摘要

被引文献

相似文献

自2005年以来,处理器设计人员已经增加了核心数量,以利用摩尔定律扩展,而不是专注于单核性能。Dennard扩展的失败(向多核部分的转变部分是对此的回应)可能很快就会限制多核扩展,就像单核扩展已经被削减一样。本文通过结合设备扩展、单核扩展和多核扩展来模拟多核扩展限制,以衡量未来五代技术中一组并行工作负载的加速潜力。对于设备缩放,我们使用ITRS预测和一组更保守的设备缩放参数。为了对单核扩展进行建模,我们结合了来自150多个处理器的联合收割机测量结果,以获得面积/性能和功耗/性能的帕累托最优边界。最后,为了对多核扩展进行建模,我们构建了一个详细的性能模型,包括上限性能和下限核心功耗。我们研究的多核设计包括单线程CPU和大规模线程GPU的多核芯片组织,具有对称,非对称,动态和组合拓扑。该研究表明,无论芯片组织和拓扑如何,多核扩展的功率限制程度都没有得到计算社区的广泛认可。即使在22 nm(仅一年后),21%的固定尺寸芯片必须断电,而在8 nm,这一数字增长到50%以上。到2024年,在常用的并行工作负载中,平均加速只有7.9倍,距离每代性能翻倍的目标还有近24倍的差距。
Since 2005, processor designers have increased core counts to exploit Moore's Law scaling, rather than focusing on single-core performance. The failure of Dennard scaling, to which the shift to multicore parts is partially a response, may soon limit multicore scaling just as single-core scaling has been curtailed. This paper models multicore scaling limits by combining device scaling, single-core scaling, and multicore scaling to measure the speedup potential for a set of parallel workloads for the next five technology generations. For device scaling, we use both the ITRS projections and a set of more conservative device scaling parameters. To model single-core scaling, we combine measurements from over 150 processors to derive Pareto-optimal frontiers for area/performance and power/performance. Finally, to model multicore scaling, we build a detailed performance model of upper-bound performance and lower-bound core power. The multicore designs we study include single-threaded CPU-like and massively threaded GPU-like multicore chip organizations with symmetric, asymmetric, dynamic, and composed topologies. The study shows that regardless of chip organization and topology, multicore scaling is power limited to a degree not widely appreciated by the computing community. Even at 22 nm (just one year from now), 21% of a fixed-size chip must be powered off, and at 8 nm, this number grows to more than 50%. Through 2024, only 7.9× average speedup is possible across commonly used parallel workloads, leaving a nearly 24-fold gap from a target of doubled performance per generation.