MegaProto: 1 TFlops/10kW Rack Is Feasible Even with Only Commodity Technology

MegaProto: 1 TFlops/10kW Rack Is Feasible Even with Only Commodity Technology
复制标题

DOI:
10.1109/sc.2005.45
复制
发表时间:
2005-11
期刊:
ACM/IEEE SC 2005 Conference (SC'05)
影响因子:
--
通讯作者:
H. Nakashima;Hiroshi Nakamura;M. Sato;T. Boku;S. Matsuoka;D. Takahashi;Y. Hotta
H. Nakashima;Hiroshi Nakamura;M. Sato;T. Boku;S. Matsuoka;D. Takahashi;Y. Hotta
中科院分区:
其他
文献类型:
--
作者:
H. Nakashima;Hiroshi Nakamura;M. Sato;T. Boku;S. Matsuoka;D. Takahashi;Y. Hotta

文献摘要

相似文献

在我们的研究项目“基于低功耗技术和并行建模的大规模计算”中,我们声称可以用密集安装的低功耗商用处理器构建百万级并行系统。“MegaProto”是一个概念验证的低功耗和高性能集群,仅使用商品组件来实现这一要求。一个机架系统由32个主板“集群单元”组成,这些主板“集群单元”具有1个U高度和商品交换机,以将它们相互连接以及与其他机架连接。每个集群单元容纳16个低功耗美元大小的商品PC架构子板,以及一个高带宽,2 Gbps的每处理器嵌入式交换网络的基础上千兆以太网。单机架系统的峰值性能在第一个版本中为0.48 TFlops,在第二个版本中将通过处理器/子板升级提高到1.02 TFlops。该系统每个机架的功耗约为10千瓦或更少,使用32 Gbps二分带宽的功率感知机架内网络,功率效率可达100 MFlops/W,而额外的2.4千瓦将使其达到足够大的256 Gbps。性能研究表明,在大多数NPB程序中,即使是第一个版本也明显优于由双功耗处理器组成的传统高端1U服务器。本文还研究了目前的自动化DVS控制如何为HPC并行程序节省功率沿着局限性。
In our research project "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we claim that a million-scale parallel system could be built with densely mounted low-power commodity processors. "MegaProto" is a proof-of-concept low-power and highperformance cluster build only with commodity components to implement this claim. A one-rack system is composed of 32 motherboard "cluster units" of 1 U-height and commodity switches to interconnect them mutually as well as with other racks. Each cluster unit houses 16 low-power dollarbill- sized commodity PC-architecture daughterboards, together with a high bandwidth, 2 Gbps per processor embedded switched network based on Gigabit Ethernet. The peak performance of a one-rack system is 0.48 TFlops for the first version and will improve to 1.02 TFlops in the second version through a processor/daughterboard upgrade. The system consumes about 10 kW or less per rack, resulting in 100 MFlops/W power efficiency with a power-aware intrarack network of 32 Gbps bisection bandwidth, while additional 2.4 kW will boost this to sufficiently large 256 Gbps. Performance studies show that even the first version significantly outperforms a conventional high-end 1U server comprised of dual power-hungry processors in a majority of NPB programs. It is also investigated how the current automated DVS control could save power for the HPC parallel programs along with its limitation.