Architecture Independent Performance Characterization and Benchmarking for Scientific Applications
Architecture Independent Performance Characterization and Benchmarking for Scientific Applications
复制标题
科学应用的独立于架构的性能表征和基准测试
DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
H. Shan
中科院分区:
文献类型:
--
作者:
Erich Strohmaier;H. Shan
Architecture Independent Performance Characterization and Benchmarking for Scientific Applications Erich Strohmaier, Hongzhang Shan Future Technology Group Lawrence Berkeley National Laboratory One Cyclotron Road, CA 94720 {estrohmaier, hshan@lbl.gov} Abstract A simple, tunable, synthetic benchmark with a per- formance directly related to applications would be of great benefit to the scientific computing community. In this paper, we present a novel approach to develop such a benchmark. The initial focus of this project is on data access performance of scientific applications. First a hardware independent characterization of code per- formance in terms of address streams is developed. The parameters chosen to characterize a single address stream are related to regularity, size, spatial, and tem- poral locality. These parameters are then used to im- plement a synthetic benchmark program that mimics the performance of a corresponding code. To test the valid- ity of our approach we performed experiments using five test kernels on six different platforms. The performance of most of our test kernels can be approximated by a single synthetic address stream. However in some cases overlapping two address streams is necessary to achieve a good approximation. Synthetic benchmarks are much easier to develop, maintain, and use. Synthetic benchmarks such as Linpack [1], NAS Parallel Benchmarks [2], PARKBENCH [3], or SPLASH2 [9] have been developed based on specific codes, which were selected ad-hoc. However, they only reflect the performance of a very narrow set of applica- tions at best and cannot serve as general benchmark against which the progress in real application perform- ance could be judged. It is therefore of great interest to develop a tunable synthetic benchmark, the performance of which we can relate to the performance of our applica- tion codes. This requires that we develop a parametric characterization of the performance behavior of scientific application codes first, which will then lead us directly to the design of such a synthetic benchmark. The initial assumption of this study is that the per- formance behavior of codes can be characterized by a small set of performance factors, which are specific to the code and independent of the computer architecture. The performance of a class of codes with similar characteris- tic performance factors should then be closely related to each other on different architectures. A synthetic bench- mark implemented such that its execution profile could be tuned by corresponding input parameters to match the chosen characteristic performance factors, could then be used as a proxy for the performance behavior of codes with similar characteristic parameters. Therefore such a benchmark can be used as a more realistic indicator of achievable performance on existing or new platforms. Since the design of the benchmark is only guided by the application characterization methodology and is inde- pendent of any specific architecture, the benchmark can be used for a long time and across all platforms. As a synthetic benchmark it greatly reduces maintenance and provides increased portability. Due its compact code size it also is substantially easier to use with simulators. During the last decades memory access became the dominant performance factor for many codes. Therefore data access is the initial focus of our application perform- 1. Introduction When benchmarking computer systems ultimately the performance of our scientific application codes is most important to us. However using them for performance studies requires an enormous amount of time and effort and studies based on them tend to have limited architec- tural scope. This makes fair comparisons across architec- tural domains cumbersome, as they require a substantial effort for code adaptation and optimization. The amount of available results therefore tends to be limited. Applica- tion codes also tend to have a very limited meaningful life-time as benchmarks, as the actual applications tend to evolve and change at a rapid pace making special benchmark versions obsolete. Recent examples for this type of benchmarking efforts include SPEC HPC [20] and The Matrix [4].