Performance Modeling and Comparative Analysis of the MILC Lattice QCD Application su3_rmd

Performance Modeling and Comparative Analysis of the MILC Lattice QCD Application su3_rmd
复制标题

MILC 晶格 QCD 应用的性能建模和比较分析 su3_rmd

DOI:
10.1109/ccgrid.2012.123
复制
发表时间:
2012
期刊:
2012 12th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (ccgrid 2012)
影响因子:
--
通讯作者:
T. Hoefler
T. Hoefler
中科院分区:
--
文献类型:
--
作者:
G. Bauer;S. Gottlieb;T. Hoefler

文献摘要

被引文献

相似文献

当HPC进入Petascale并为Exascale做准备时,应用程序性能建模是应用和系统开发的重要组成部分。但是,由于测量和噪声效应的自然变化,平行系统的性能建模是一项艰巨的任务。在本文中,我们为在多种并行计算平台上的晶状体量子量子Chromo动力学字段应用于无处不在的HPC应用程序SU3 RMD的半分析性能模型方法SU3 RMD提供了详细的示例。我们应用在自然科学中众所周知的统计技术来对输入系统的方差进行建模。使用一个简单的分析模型来捕获代码的主要特征,例如传递消息的数字和大小以及串行代码块的调用计数以及统计上声音曲线拟合方法,我们开发了一个准确的性能模型,并使用它来表征应用程序各种目标体系结构的性能。我们的拟合技术使我们能够表征给定系统上不同性能观察的方差,并显示来自不同来源的噪声的影响。我们开发的技术可以应用于一系列批量同步应用。以这个详细的例子,我们旨在激励科学计算社区开发和使用类似的性能模型进行软件开发和维护。
Application performance modeling is an essential part of application and system development as HPC moves into the petascale and prepares for the exascale. However, performance modeling of parallel systems is a difficult task due to natural variations in measurements and noise effects. In this paper, we give a detailed example for a semi-analytical performance-modeling method applied to the ubiquitous HPC application su3 rmd from the lattice Quantum Chromo dynamics field on a variety of parallel computing platforms. We apply statistical techniques that are well known in natural sciences to model the variance in the input system. Using a simple analytical model to capture the main characteristics of the code, such as numbers and sizes of passed messages and invocation counts of serial code blocks in conjunction with statistically sound curve fitting methods, we develop an accurate performance model and use it to characterize application performance on various target architectures. Our fitting techniques allow us to characterize the variance of different performance observations on a given system and show the influence of noise from different sources. The techniques we developed can be applied to a wide class of bulk-synchronous applications. With this detailed example, we aim to motivate the scientific computing community to develop and use similar performance models for software development and maintenance.