Cache Optimization for Coarse Grain Task Parallel Processing Using Inter-Array Padding

Cache Optimization for Coarse Grain Task Parallel Processing Using Inter-Array Padding
复制标题

使用阵列间填充的粗粒度任务并行处理的缓存优化

DOI:
10.1007/978-3-540-24644-2_5
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
H. Kasahara
H. Kasahara
中科院分区:
--
文献类型:
--
作者:
K. Ishizaka;M. Obata;H. Kasahara

文献摘要

被引文献

相似文献

多处理器系统的广泛应用使得自动并行编译器变得更加重要。为了通过编译器进一步提高多处理器系统的性能,多粒并行化是非常重要的。在多粒度并行化中,除了传统的循环并行外,还采用了循环和子程序之间的粗粒度任务并行和语句之间的近细粒度并行。此外,有效使用缓存的局部性优化对性能提高也很重要。为了减少宏任务之间的缓存冲突缺失,本文采用数据定位方案,将共享相同数组的循环分解为适合缓存大小的循环,并在同一处理器上连续执行分解后的循环。在Sun Ultra 80(4pe)上的性能评估中,实现该方案的OSCAR编译器在SPEC CFP95 tomcatv、swim hydro2d和turb3d程序的平均性能上比Sun Forte编译器自动循环并行化的最大性能提高了2.5倍。OSCAR编译器在IBM RS/6000 44p-270(4pe)上的速度比XLF编译器快2.1倍。
The wide use of multiprocessor system has been making automatic parallelizing compilers more important. To improve the performance of multiprocessor system more by compiler, multigrain parallelization is important. In multigrain parallelization, coarse grain task parallelism among loops and subroutines and near fine grain parallelism among statements are used in addition to the traditional loop parallelism. In addition, locality optimization to use cache effectively is also important for the performance improvement. This paper describes inter-array padding to minimize cache conflict misses among macro-tasks with data localization scheme which decomposes loops sharing the same arrays to fit cache size and executes the decomposed loops consecutively on the same processor. In the performance evaluation on Sun Ultra 80(4pe), OSCAR compiler on which the proposed scheme is implemented gave us 2.5 times speedup against the maximum performance of Sun Forte compiler automatic loop parallelization at the average of SPEC CFP95 tomcatv, swim hydro2d and turb3d programs. Also, OSCAR compiler showed 2.1 times speedup on IBM RS/6000 44p-270(4pe) against XLF compiler.