Identifying the Culprits Behind Network Congestion

Identifying the Culprits Behind Network Congestion
复制标题

DOI:
10.1109/ipdps.2015.92
复制
发表时间:
2015-05
期刊:
2015 IEEE International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
A. Bhatele;Andrew R. Titus;Jayaraman J. Thiagarajan;Nikhil Jain;T. Gamblin;P. Bremer;M. Schulz;L. Kalé
A. Bhatele;Andrew R. Titus;Jayaraman J. Thiagarajan;Nikhil Jain;T. Gamblin;P. Bremer;M. Schulz;L. Kalé
中科院分区:
其他
文献类型:
--
作者:
A. Bhatele;Andrew R. Titus;Jayaraman J. Thiagarajan;Nikhil Jain;T. Gamblin;P. Bremer;M. Schulz;L. Kalé

文献摘要

被引文献

相似文献

网络拥塞是导致高通信量并行应用性能下降、性能变化和可扩展性差的主要原因之一。然而,现代互联网络上的网络拥塞的原因和机制还没有得到很好的理解。我们需要新的方法来分析,建模和预测这一关键行为,以提高大规模并行应用程序的性能。本文应用监督学习算法,如极端随机树森林和梯度提升回归树,对通信数据和应用程序执行时间进行回归分析。使用来自多个执行的数据,我们创建模型来预测通信繁重的并行应用程序的执行时间。该分析还确定了对网络拥塞和内部执行时间影响最大的功能和相关硬件组件。本文提出的思想具有广泛的适用性:预测不同数量的节点,或不同的输入数据集,甚至是未知代码的执行时间,确定应用程序的最佳配置参数,并找到不同架构上网络拥塞的根本原因。
Network congestion is one of the primary causes of performance degradation, performance variability and poor scaling in communication-heavy parallel applications. However, the causes and mechanisms of network congestion on modern interconnection networks are not well understood. We need new approaches to analyze, model and predict this critical behaviour in order to improve the performance of large-scale parallel applications. This paper applies supervised learning algorithms, such as forests of extremely randomized trees and gradient boosted regression trees, to perform regression analysis on communication data and application execution time. Using data derived from multiple executions, we create models to predict the execution time of communication-heavy parallel applications. This analysis also identifies the features and associated hardware components that have the most impact on network congestion and intern, on execution time. The ideas presented in this paper have wide applicability: predicting the execution time on a different number of nodes, or different input datasets, or even for an unknown code, identifying the best configuration parameters for an application, and finding the root causes of network congestion on different architectures.