Performance Prediction for Large-Scale Parallel Applications Using Representative Replay

Performance Prediction for Large-Scale Parallel Applications Using Representative Replay
复制标题

使用代表性重放的大规模并行应用程序的性能预测

DOI:
10.1109/tc.2015.2479630
复制
发表时间:
2016-07
影响因子:
3.7
通讯作者:
Li Keqin
Li Keqin
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhai Jidong;Chen Wenguang;Zheng Weimin;Li Keqin

文献摘要

参考文献

被引文献

相似文献

自动预测并行应用程序的性能一直是高性能计算领域的一个长期目标。然而,准确的性能预测是具有挑战性的,因为并行应用程序的执行时间是由几个因素决定的,例如顺序计算时间、通信时间及其复杂的相互作用。尽管前人已经做了很多努力,但在大规模并行应用中,准确估计每个进程的顺序计算时间仍然是一个有待解决的问题。在本文中,我们提出了一种新的方法来获取精确的顺序计算时间,使用并行调试技术称为确定性重播。这种方法的主要优点是,我们只需要目标平台的单个节点,而不需要整个目标平台可用。因此,使用这种方法,我们可以简单地测量每个进程在目标节点上的实际顺序计算时间。此外,我们观察到,在并行应用中,不仅在每个进程内,而且在不同进程之间,都有很大的计算相似性。基于这一观察,我们进一步提出可以显著减少重播开销的代表性重播,因为我们只需要重播代表性流程的部分迭代,而不是所有的迭代。最后,我们实现了一个完整的性能预测系统,称为Phantom,它结合了上述计算时间获取方法和跟踪驱动模拟器。我们在传统的HPC平台和最新的Amazon EC2云平台上验证了我们的方法。在这两种类型的平台上,我们的方法的预测误差平均小于7%,最多2500个过程。
Automatically predicting performance of parallel applications has been a long-standing goal in the area of high performance computing. However, accurate performance prediction is challenging, since the execution time of parallel applications is determined by several factors, such as sequential computation time, communication time and their complex interactions. Despite previous efforts, accurately estimating the sequential computation time in each process for large-scale parallel applications remains an open problem. In this paper, we propose a novel approach to acquiring accurate sequential computation time using a parallel debugging technique called deterministic replay. The main advantage of our approach is that we only need a single node of a target platform but the whole target platform does not need to be available. Therefore, with this approach we can simply measure the real sequential computation time on a target node for each process on by one. Moreover, we observe that there is great computation similarity in parallel applications, not only within each process but also among different processes. Based on this observation, we further propose representative replay that can significantly reduce replay overhead, because we only need to replay partial iterations for representative processes instead of all of them. Finally, we implement a complete performance prediction system, called Phantom, which combines the above computation-time acquisition approach and a trace-driven simulator. We validate our approach on both traditional HPC platforms and the latest Amazon EC2 cloud platform. On both types of platforms, prediction error of our approach is less than 7 percent on average up to 2,500 processes.
DOI: 10.1006/jpdc.1997.1346
发表时间: 1997
期刊: J. Parallel Distributed Comput.
影响因子: --
作者:
Albert D. Alexandrov;M. Ionescu;K. Schauser;C. Scheiman
通讯作者: Albert D. Alexandrov;M. Ionescu;K. Schauser;C. Scheiman
DOI: 10.1007/978-3-540-75416-9_41
发表时间: 2007-09
期刊: --
影响因子: --
作者:
A. Bouteiller;G. Bosilca;J. Dongarra
通讯作者: A. Bouteiller;G. Bosilca;J. Dongarra
DOI: --
发表时间: 2002
期刊: --
影响因子: --
作者:
A. Snavely;L. Carrington;N. Wolter;J. Labarta;R. Badia;A. Purkayastha
通讯作者: A. Snavely;L. Carrington;N. Wolter;J. Labarta;R. Badia;A. Purkayastha
DOI: 10.1145/2063384.2063451
发表时间: 2011-11
期刊: 2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子: --
作者:
I. Laguna;T. Gamblin;B. Supinski;S. Bagchi;G. Bronevetsky;D. Ahn;M. Schulz;B. Rountree
通讯作者: I. Laguna;T. Gamblin;B. Supinski;S. Bagchi;G. Bronevetsky;D. Ahn;M. Schulz;B. Rountree
DOI: 10.1145/1504176.1504213
发表时间: 2009-02
期刊: --
影响因子: --
作者:
Ruini Xue;Xuezheng Liu;Ming Wu;Zhenyu Guo;Wenguang Chen;Weimin Zheng;Zheng Zhang;G. Voelker
通讯作者: Ruini Xue;Xuezheng Liu;Ming Wu;Zhenyu Guo;Wenguang Chen;Weimin Zheng;Zheng Zhang;G. Voelker