Dynamic pointer alignment: tiling and communication optimizations for parallel pointer-based computations

Dynamic pointer alignment: tiling and communication optimizations for parallel pointer-based computations
复制标题

动态指针对齐:基于指针的并行计算的平铺和通信优化

DOI:
--
复制
发表时间:
1997
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
A. Chien
A. Chien
中科院分区:
--
文献类型:
--
作者:
Xingbin Zhang;A. Chien

文献摘要

被引文献

相似文献

循环平铺和通信优化(如消息流水线和聚合)可以通过主动管理存储和数据移动来实现优化和强大的内存性能。在本文中,我们将这些技术推广到基于指针的数据结构(PBDS)。我们的方法,动态指针对齐(DPA),有两个组成部分。编译器将程序分解为非阻塞线程,这些线程在特定的指针上操作,并使用相应的指针标记线程创建站点。在运行时,从指针到依赖线程的显式映射在线程创建时更新,并用于动态调度线程和通信,使得使用相同对象的线程一起执行,通信与本地工作重叠,并且消息被聚合。我们已经实现了DPA,以优化对并行机上全局PBDS的远程读取。我们对使用复杂PBDS的两个应用程序(Barnes-Hut和FMM)的力计算阶段的经验结果表明,DPA通过在CRAY T3 D上实现平铺和通信优化,实现了良好的绝对性能和加速。
Loop tiling and communication optimization, such as message pipelining and aggregation, can achieve optimized and robust memory performance by proactively managing storage and data movement. In this paper, we generalize these techniques to pointer-based data structures (PBDSs). Our approach, dynamic pointer alignment (DPA), has two components. The compiler decomposes a program into non-blocking threads that operate on specific pointers and labels thread creation sites with their corresponding pointers. At runtime, an explicit mapping from pointers to dependent threads is updated at thread creation and is used to dynamically schedule both threads and communication, such that threads using the same objects execute together, communication overlaps with local work, and messages are aggregated. We have implemented DPA to optimize remote reads to global PBDSs on parallel machines. Our empirical results on the force computation phases of two applications that use sophisticated PBDSs, Barnes-Hut and FMM, show that DPA achieves good absolute performance and speedups by enabling tiling and communication optimization on the CRAY T3D.