Active Measurement of the Impact of Network Switch Utilization on Application Performance

Active Measurement of the Impact of Network Switch Utilization on Application Performance
复制标题

主动测量网络交换机利用率对应用程序性能的影响

DOI:
10.1109/ipdps.2014.28
复制
发表时间:
2014
期刊:
2014 IEEE 28th International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
G. Bronevetsky
G. Bronevetsky
中科院分区:
--
文献类型:
--
作者:
Marc Casas;G. Bronevetsky

文献摘要

被引文献

相似文献

节点间网络是高性能计算(HPC)系统的一项关键功能,它将HPC系统与能力较低的机器类别区分开来。然而,尽管HPC计算节点的性能非常高,但其计算能力的不断提高以及应用程序通信需求的相关增长使网络性能成为常见的性能瓶颈。为了在网络限制的情况下实现高性能,应用程序开发人员需要工具来测量其应用程序的网络利用率,并告知他们网络的通信容量如何与其应用程序的性能相关。本文提出了一种新的性能测量和分析方法的基础上的经验测量的网络行为。我们的方法使用两个基准测试,注入额外的网络通信。第一种方法探测软件组件(应用程序或单个任务)使用的网络部分,以确定网络争用的存在和严重性。第二种方法是在软件组件运行以评估其在能力较低的网络上的性能时,或者在它与其他软件组件共享网络时,积极地注入网络流量。然后,我们结合联合收割机的信息,从两种类型的实验,以预测多个软件组件(例如,一个MPI应用程序的多个进程)时,他们共享一个单一的网络所经历的性能放缓。我们的方法适用于个人网络交换机,并证明了6个代表性的HPC应用程序和预测的36个可能的应用程序对的性能下降。我们预测的平均误差小于10%。
Inter-node networks are a key capability of High-Performance Computing (HPC) systems that differentiates them from less capable classes of machines. However, in spite of their very high performance, the increasing computational power of HPC compute nodes and the associated rise in application communication needs make network performance a common performance bottleneck. To achieve high performance in spite of network limitations application developers require tools to measure their applications' network utilization and inform them about how the network's communication capacity relates to the performance of their applications. This paper presents a new performance measurement and analysis methodology based on empirical measurements of network behavior. Our approach uses two benchmarks that inject extra network communication. The first probes the fraction of the network that is utilized by a software component (an application or an individual task) to determine the existence and severity of network contention. The second aggressively injects network traffic while a software component runs to evaluate its performance on less capable networks or when it shares the network with other software components. We then combine the information from the two types of experiments to predict the performance slowdown experienced by multiple software components (e.g. multiple processes of a single MPI application) when they share a single network. Our methodology is applied to individual network switches and demonstrated taking 6 representative HPC applications and predicting the performance slowdowns of the 36 possible application pairs. The average error of our predictions is less than 10%.