Cray Cascade: A scalable HPC system based on a Dragonfly network

Cray Cascade: A scalable HPC system based on a Dragonfly network
复制标题

DOI:
10.1109/sc.2012.39
复制
发表时间:
2012-11
期刊:
2012 International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Greg Faanes;A. Bataineh;D. Roweth;T. Court;E. Froese;Robert Alverson;Tim Johnson;Joe Kopnick;
Greg Faanes;A. Bataineh;D. Roweth;T. Court;E. Froese;Robert Alverson;Tim Johnson;Joe Kopnick;
中科院分区:
其他
文献类型:
--
作者:
Greg Faanes;A. Bataineh;D. Roweth;T. Court;E. Froese;Robert Alverson;Tim Johnson;Joe Kopnick;

文献摘要

被引文献

相似文献

对于许多应用来说,更高的全局带宽需求和更低的网络成本促使高性能计算系统使用Dragonfly网络拓扑。本文提出了基于Dragonfly[1]网络拓扑的分布式存储系统Cray Cascade的体系结构。我们描述了系统的结构、蜻蜓网络和路由算法。我们描述了一组支持主流高性能计算应用和新兴全局地址空间编程模型的高级特性。我们结合了原型系统的性能结果和大型系统的仿真数据。我们展示了蜻蜓拓扑的价值,以及通过广泛使用自适应路由获得的好处。
Higher global bandwidth requirement for many applications and lower network cost have motivated the use of the Dragonfly network topology for high performance computing systems. In this paper we present the architecture of the Cray Cascade system, a distributed memory system based on the Dragonfly [1] network topology. We describe the structure of the system, its Dragonfly network and the routing algorithms. We describe a set of advanced features supporting both mainstream high performance computing applications and emerging global address space programing models. We present a combination of performance results from prototype systems and simulation data for large systems. We demonstrate the value of the Dragonfly topology and the benefits obtained through extensive use of adaptive routing.