Maximizing Throughput on a Dragonfly Network

Maximizing Throughput on a Dragonfly Network
复制标题

最大化 Dragonfly 网络的吞吐量

DOI:
10.1109/sc.2014.33
复制
发表时间:
2014
期刊:
SC14: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
L. Kalé
L. Kalé
中科院分区:
--
文献类型:
--
作者:
Nikhil Jain;A. Bhatele;Xiang Ni;N. Wright;L. Kalé

文献摘要

被引文献

相似文献

互连网络是大型超级计算机的关键资源。网络拓扑结构,它提供了一个低的网络直径和大的二分带宽,正在探索作为一个有前途的选择,建立多Petaflop的和Exaflop的系统。与广泛研究的环面网络,最好的选择的消息路由和工作布局的策略,为turbine拓扑结构没有很好地理解。本文的目的是分析行为的机器构建使用一个可扩展网络的各种路由策略,就业安置政策,和应用程序的通信模式。我们的研究是基于一种新的模型,预测直接,间接和自适应路由策略的各个环节上的流量。我们分析了个别通信模式和一些常见的并行作业工作负载的结果。本文中提出的预测是针对具有92,160个高基数路由器和880万个核心的100+ Petaflop的原型机器。
Interconnection networks are a critical resource for large supercomputers. The dragonfly topology, which provides a low network diameter and large bisection bandwidth, is being explored as a promising option for building multi-Petaflop's and Exaflop's systems. Unlike the extensively studied torus networks, the best choices of message routing and job placement strategies for the dragonfly topology are not well understood. This paper aims at analyzing the behavior of a machine built using a dragonfly network for various routing strategies, job placement policies, and application communication patterns. Our study is based on a novel model that predicts traffic on individual links for direct, indirect, and adaptive routing strategies. We analyze results for individual communication patterns and some common parallel job workloads. The predictions presented in this paper are for a 100+ Petaflop's prototype machine with 92,160 high radix routers and 8.8 million cores.