Scaling the bandwidth wall: challenges in and avenues for CMP scaling

Scaling the bandwidth wall: challenges in and avenues for CMP scaling
复制标题

DOI:
10.1145/1555754.1555801
复制
发表时间:
2009-06
期刊:
--
影响因子:
--
通讯作者:
Brian Rogers;A. Krishna;Gordon B. Bell;K. V. Vu;Xiaowei Jiang;Yan Solihin
Brian Rogers;A. Krishna;Gordon B. Bell;K. V. Vu;Xiaowei Jiang;Yan Solihin
中科院分区:
其他
文献类型:
--
作者:
Brian Rogers;A. Krishna;Gordon B. Bell;K. V. Vu;Xiaowei Jiang;Yan Solihin

文献摘要

被引文献

相似文献

随着晶体管密度按照摩尔定律继续以指数速度增长,许多芯片多处理器 (CMP) 系统的目标是按比例扩展片上内核的数量。不幸的是,与内核数量的预期增长相比,片外存储器带宽容量预计增长缓慢。这就造成了这样一种情况,即每个内核可用于从片外存储器加载数据的片外带宽量会减少。片外带宽成为性能和吞吐量瓶颈的情况被称为带宽墙问题。在本研究中,我们试图回答两个问题:(1)带宽墙问题在多大程度上限制了未来的多核扩展,以及(2)各种带宽节约技术能够在多大程度上缓解这个问题。为了解决这些问题,我们开发了一个简单但功能强大的分析模型来预测 CMP 在内存流量容量增长有限的情况下可以支持的片上内核数量。我们发现带宽墙会严重限制核心扩展。当从平衡的 8 核 CMP 开始时,在四代技术中,核心数量只能扩展到 24 个,而不是按比例扩展时的 128 个核心,而不会增加内存流量要求。我们发现,我们评估的各种单独的带宽节省技术对核心扩展具有广泛的影响,并且当组合在一起时,这些技术有可能实现最多 4 代技术的超比例核心扩展。
As transistor density continues to grow at an exponential rate in accordance to Moore's law, the goal for many Chip Multi-Processor (CMP) systems is to scale the number of on-chip cores proportionally. Unfortunately, off-chip memory bandwidth capacity is projected to grow slowly compared to the desired growth in the number of cores. This creates a situation in which each core will have a decreasing amount of off-chip bandwidth that it can use to load its data from off-chip memory. The situation in which off-chip bandwidth is becoming a performance and throughput bottleneck is referred to as the bandwidth wall problem. In this study, we seek to answer two questions: (1) to what extent does the bandwidth wall problem restrict future multicore scaling, and (2) to what extent are various bandwidth conservation techniques able to mitigate this problem. To address them, we develop a simple but powerful analytical model to predict the number of on-chip cores that a CMP can support given a limited growth in memory traffic capacity. We find that the bandwidth wall can severely limit core scaling. When starting with a balanced 8-core CMP, in four technology generations the number of cores can only scale to 24, as opposed to 128 cores under proportional scaling, without increasing the memory traffic requirement. We find that various individual bandwidth conservation techniques we evaluate have a wide ranging impact on core scaling, and when combined together, these techniques have the potential to enable super-proportional core scaling for up to 4 technology generations.