Towards scalable, energy-efficient, bus-based on-chip networks

Towards scalable, energy-efficient, bus-based on-chip networks
复制标题

DOI:
10.1109/hpca.2010.5416639
复制
发表时间:
2010-04
期刊:
HPCA - 16 2010 The Sixteenth International Symposium on High-Performance Computer Architecture
影响因子:
--
通讯作者:
Aniruddha N. Udipi;N. Muralimanohar;R. Balasubramonian
Aniruddha N. Udipi;N. Muralimanohar;R. Balasubramonian
中科院分区:
其他
文献类型:
--
作者:
Aniruddha N. Udipi;N. Muralimanohar;R. Balasubramonian

文献摘要

被引文献

相似文献

预计未来用于众核处理器的片上网络将在能量、延迟、复杂性、验证工作和面积方面施加巨大的开销。人们普遍认为,未来应用所需的带宽只能通过采用具有复杂路由器和可扩展的基于目录的一致性协议的分组交换网络来提供。我们认为,这样的计划可能是矫枉过正,在一个精心设计的系统,除了昂贵的功率方面,因为大量的耗电路由器。我们发现,基于总线的网络与监听协议可以显着降低能耗,简化网络/协议的设计和验证,而不会损失性能。我们通过将芯片分成多个部分来实现这些特性,每个部分都有自己的广播总线,这些总线通过中央总线进一步连接。这有助于消除昂贵的路由器,但会受到长电线的能源开销的影响。我们建议使用多个布隆过滤器,以有效地跟踪数据存在于该高速缓存和限制总线广播的一个子集的段,显着降低能耗。我们进一步表明,使用操作系统页面着色有助于最大限度地提高局部性,提高了布隆过滤器的有效性。我们亦采用低摆幅布线,以进一步减少链路的能源开支。通过更多地利用片上丰富的金属预算,并采用多个地址交错总线而不是多个路由器,还可以以相对较低的成本提高性能。因此,通过结合上述所有创新,我们扩展了总线的可扩展性,并相信总线可以成为未来片上网络的可行且有吸引力的选择。与许多最先进的分组交换网络相比,我们的平均能耗降低了31倍。
It is expected that future on-chip networks for many-core processors will impose huge overheads in terms of energy, delay, complexity, verification effort, and area. There is a common belief that the bandwidth necessary for future applications can only be provided by employing packet-switched networks with complex routers and a scalable directory-based coherence protocol. We posit that such a scheme might likely be overkill in a well designed system in addition to being expensive in terms of power because of a large number of power-hungry routers. We show that bus-based networks with snooping protocols can significantly lower energy consumption and simplify network/protocol design and verification, with no loss in performance. We achieve these characteristics by dividing the chip into multiple segments, each having its own broadcast bus, with these buses further connected by a central bus. This helps eliminate expensive routers, but suffers from the energy overhead of long wires. We propose the use of multiple Bloom filters to effectively track data presence in the cache and restrict bus broadcasts to a subset of segments, significantly reducing energy consumption. We further show that the use of OS page coloring helps maximize locality and improves the effectiveness of the Bloom filters. We also employ low-swing wiring to further reduce the energy overheads of the links. Performance can also be improved at relatively low costs by utilizing more of the abundant metal budgets on-chip and employing multiple address-interleaved buses rather than multiple routers. Thus, with the combination of all the above innovations, we extend the scalability of buses and believe that buses can be a viable and attractive option for future on-chip networks. We show energy reductions of up to 31X on average compared to many state-of-the-art packet switched networks.