Architecture Independent Performance Characterization and Benchmarking for Scientific Applications

Architecture Independent Performance Characterization and Benchmarking for Scientific Applications
复制标题

科学应用的独立于架构的性能表征和基准测试

DOI:
--
复制
发表时间:
2004
期刊:
IEEE/ACM International Symposium on Modeling, Analysis, and Simulation On Computer and Telecommunication Systems
影响因子:
--
通讯作者:
H. Shan
H. Shan
中科院分区:
--
文献类型:
--
作者:
Erich Strohmaier;H. Shan

文献摘要

被引文献

相似文献

Erich Strohmaier,Hongzhang Shan Future Technology Group劳伦斯伯克利国家实验室One Cyclotron Road,CA 94720 {estrohmaier,hshan@lbl.gov}摘要一个简单的、可调的、综合的基准测试,其性能与应用程序直接相关,这将对科学计算社区有很大的好处。在本文中,我们提出了一种新的方法来开发这样的基准。该项目最初的重点是科学应用程序的数据访问性能。首先,开发了一种与硬件无关的、以地址流为单位的代码性能表征方法。选择表征单个地址流的参数与规则性、大小、空间和时间局部性有关。然后,这些参数被用来实现一个合成的基准程序,模仿相应的代码的性能。为了测试我们的方法的有效性,我们在六个不同的平台上使用五个测试内核进行了实验。我们的大多数测试内核的性能可以通过单个合成地址流来近似。然而,在某些情况下,重叠两个地址流是必要的,以实现良好的近似。合成基准测试更容易开发、维护和使用。Linpack [1]、NAS Parallel Benchmarks [2]、PARKBENCH [3]或SPLASH 2 [9]等综合基准测试都是基于特定代码开发的,这些代码都是专门选择的。然而,它们最多只能反映一组非常狭窄的应用程序的性能,不能作为判断真实的应用程序性能进展的一般基准。因此,开发一个可调的综合基准测试是非常有意义的,我们可以将其性能与我们的应用程序代码的性能联系起来。这就要求我们首先对科学应用程序代码的性能行为进行参数化描述,然后直接设计这样一个综合基准。本研究的初始假设是,代码的性能行为可以由一小组性能因素来表征,这些性能因素特定于代码并且与计算机体系结构无关。具有相似特性性能因子的一类码的性能在不同的体系结构上应该是密切相关的。合成基准被实现为使得其执行简档可以通过对应的输入参数来调整以匹配所选择的特性性能因素,然后可以用作具有类似特性参数的代码的性能行为的代理。因此,这样的基准可以用作现有或新平台上可实现的性能的更现实的指标。由于基准测试的设计仅受应用程序表征方法的指导,并且独立于任何特定的体系结构,因此基准测试可以在所有平台上长期使用。作为一个综合基准,它大大减少了维护,并提供了更高的可移植性。由于其紧凑的代码大小,它也更容易与模拟器一起使用。在过去的几十年中,存储器访问成为许多代码的主要性能因素。因此,数据访问是我们的应用程序perform- 1的初始焦点。当基准计算机系统最终我们的科学应用代码的性能是最重要的是我们。然而,使用它们进行性能研究需要大量的时间和精力,并且基于它们的研究往往具有有限的架构范围。这使得跨架构域的公平比较变得很麻烦,因为它们需要大量的代码适应和优化工作。因此,现有成果的数量往往有限。应用程序代码作为基准的有效寿命也往往非常有限,因为实际的应用程序往往以快速的速度发展和变化,使得特殊的基准版本过时。这种类型的基准测试工作的最新例子包括SPEC HPC [20]和The Matrix [4]。
Architecture Independent Performance Characterization and Benchmarking for Scientific Applications Erich Strohmaier, Hongzhang Shan Future Technology Group Lawrence Berkeley National Laboratory One Cyclotron Road, CA 94720 {estrohmaier, hshan@lbl.gov} Abstract A simple, tunable, synthetic benchmark with a per- formance directly related to applications would be of great benefit to the scientific computing community. In this paper, we present a novel approach to develop such a benchmark. The initial focus of this project is on data access performance of scientific applications. First a hardware independent characterization of code per- formance in terms of address streams is developed. The parameters chosen to characterize a single address stream are related to regularity, size, spatial, and tem- poral locality. These parameters are then used to im- plement a synthetic benchmark program that mimics the performance of a corresponding code. To test the valid- ity of our approach we performed experiments using five test kernels on six different platforms. The performance of most of our test kernels can be approximated by a single synthetic address stream. However in some cases overlapping two address streams is necessary to achieve a good approximation. Synthetic benchmarks are much easier to develop, maintain, and use. Synthetic benchmarks such as Linpack [1], NAS Parallel Benchmarks [2], PARKBENCH [3], or SPLASH2 [9] have been developed based on specific codes, which were selected ad-hoc. However, they only reflect the performance of a very narrow set of applica- tions at best and cannot serve as general benchmark against which the progress in real application perform- ance could be judged. It is therefore of great interest to develop a tunable synthetic benchmark, the performance of which we can relate to the performance of our applica- tion codes. This requires that we develop a parametric characterization of the performance behavior of scientific application codes first, which will then lead us directly to the design of such a synthetic benchmark. The initial assumption of this study is that the per- formance behavior of codes can be characterized by a small set of performance factors, which are specific to the code and independent of the computer architecture. The performance of a class of codes with similar characteris- tic performance factors should then be closely related to each other on different architectures. A synthetic bench- mark implemented such that its execution profile could be tuned by corresponding input parameters to match the chosen characteristic performance factors, could then be used as a proxy for the performance behavior of codes with similar characteristic parameters. Therefore such a benchmark can be used as a more realistic indicator of achievable performance on existing or new platforms. Since the design of the benchmark is only guided by the application characterization methodology and is inde- pendent of any specific architecture, the benchmark can be used for a long time and across all platforms. As a synthetic benchmark it greatly reduces maintenance and provides increased portability. Due its compact code size it also is substantially easier to use with simulators. During the last decades memory access became the dominant performance factor for many codes. Therefore data access is the initial focus of our application perform- 1. Introduction When benchmarking computer systems ultimately the performance of our scientific application codes is most important to us. However using them for performance studies requires an enormous amount of time and effort and studies based on them tend to have limited architec- tural scope. This makes fair comparisons across architec- tural domains cumbersome, as they require a substantial effort for code adaptation and optimization. The amount of available results therefore tends to be limited. Applica- tion codes also tend to have a very limited meaningful life-time as benchmarks, as the actual applications tend to evolve and change at a rapid pace making special benchmark versions obsolete. Recent examples for this type of benchmarking efforts include SPEC HPC [20] and The Matrix [4].