Compiling affine loop nests for distributed-memory parallel architectures

Compiling affine loop nests for distributed-memory parallel architectures
复制标题

为分布式内存并行架构编译仿射循环嵌套

DOI:
10.1145/2503210.2503289
复制
发表时间:
2013
期刊:
2013 SC - International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
Uday Bondhugula
Uday Bondhugula
中科院分区:
--
文献类型:
--
作者:
Uday Bondhugula

文献摘要

被引文献

相似文献

针对分布式存储并行体系结构,提出了一种编译具有仿射依赖关系的任意嵌套循环的新技术。我们的框架被实现为使用多面体模型的源代码级转换器,并生成使用消息传递接口(MPI)库表示的通信的并行代码。与所有以前的方法相比,我们的方法要么是(1)关于所处理的输入代码的通用性,要么是(2)通信代码的效率,或者两者兼而有之。我们给出了一个多核集群上的实验结果,证明了它的有效性。在某些情况下,我们生成的代码的性能优于手动并行化的代码,而在另一种情况下,我们生成的代码的性能低于手动并行代码的25%。据我们所知,这是第一个报告输入程序和转换技术的端到端全自动分布式内存并行化和代码生成的工作,就像我们所允许的那样。
We present new techniques for compilation of arbitrarily nested loops with affine dependences for distributed-memory parallel architectures. Our framework is implemented as a source-level transformer that uses the polyhedral model, and generates parallel code with communication expressed with the Message Passing Interface (MPI) library. Compared to all previous approaches, ours is a significant advance either (1) with respect to the generality of input code handled, or (2) efficiency of communication code, or both. We provide experimental results on a cluster of multicores demonstrating its effectiveness. In some cases, code we generate outperforms manually parallelized codes, and in another case is within 25% of it. To the best of our knowledge, this is the first work reporting end-to-end fully automatic distributed-memory parallelization and code generation for input programs and transformation techniques as general as those we allow.