PIPES: A Language and Compiler for Task-Based Programming on Distributed-Memory Clusters

PIPES: A Language and Compiler for Task-Based Programming on Distributed-Memory Clusters
复制标题

PIPES:分布式内存集群上基于任务的编程语言和编译器

DOI:
--
复制
发表时间:
2016
期刊:
International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Vivek Sarkar
Vivek Sarkar
中科院分区:
--
文献类型:
--
作者:
Martin Kong;L. Pouchet;P. Sadayappan;Vivek Sarkar

文献摘要

被引文献

相似文献

在共享内存计算机集群上运行的应用程序通常使用OpenMP+MPI实现。使用基于任务的编程可以极大地提高生产率,这是一种用户表达任务之间的数据和控制流关系的范例,为运行时提供了放置和调度任务的最大自由。虽然生产率提高了,但高性能执行仍然具有挑战性:并行算法的实施通常需要特定的任务布局和通信策略,以减少节点间通信并利用数据局部性。在这项工作中,我们提出了一个新的宏数据流编程环境,用于分布式内存集群,基于Intel并发集合(CNC)运行时。我们的语言扩展允许用户定义虚拟拓扑、任务映射、以任务为中心的数据放置、任务和通信调度等。我们引入了一个编译器来自动生成英特尔CNC C++运行时,并进行了关键的自动优化,包括任务粗化和合并。我们在各种科学计算上对我们的方法进行了实验验证,证明了我们的方法的效率和性能。
Applications running on clusters of shared-memory computers are often implemented using OpenMP+MPI. Productivity can be vastly improved using task-based programming, a paradigm where the user expresses the data and control-flow relations between tasks, offering the runtime maximal freedom to place and schedule tasks. While productivity is increased, high-performance execution remains challenging: the implementation of parallel algorithms typically requires specific task placement and communication strategies to reduce internode communications and exploit data locality. In this work, we present a new macro-dataflow programming environment for distributed-memory clusters, based on the Intel Concurrent Collections (CnC) runtime. Our language extensions let the user define virtual topologies, task mappings, task-centric data placement, task and communication scheduling, etc. We introduce a compiler to automatically generate Intel CnC C++ run-time, with key automatic optimizations including task coarsening and coalescing. We experimentally validate our approach on a variety of scientific computations, demonstrating both productivity and performance.