Tapas: An Implicitly Parallel Programming Framework for Hierarchical N-Body Algorithms

Tapas: An Implicitly Parallel Programming Framework for Hierarchical N-Body Algorithms
复制标题

Tapas:分层 N 体算法的隐式并行编程框架

DOI:
10.1109/icpads.2016.0145
复制
发表时间:
2016
期刊:
2016 IEEE 22nd International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
通讯作者:
and Satoshi Matsuoka
and Satoshi Matsuoka
中科院分区:
--
文献类型:
--
作者:
Keisuke Fukuda;Motohiko Matsuda;Naoya Maruyama;Rio Yokota;Kenjiro Taura;and Satoshi Matsuoka

文献摘要

相似文献

Tapas是我们新的C++编程框架,用于大规模异构超级计算机上的分层算法,如N-body。虽然N体及其变体在科学应用中广泛使用,但它们的正确实现通常很难在这样的现代机器上实现,因为算法是不规则的,复杂的,并且涉及分布式节点上的显式任务并行编程。由于大规模分布式内存上的不规则数据访问,在库或框架中封装复杂性一直是一个挑战。Tapas通过使用C++模板元编程将用户干净的隐式并行程序转换为异构多核多节点环境上的检查器-执行器风格的代码来解决这个问题。Tapas上的快速多极子方法的原型实现展示了与ExaFMM相当的性能和扩展性,ExaFMM是FMM最快的手动调优实现,以及数百个GPU的有效使用。具体而言,串行性能是ExaFMM的95%,而使用多达1500个CPU内核的分布式内存强扩展评估显示ExaFMM性能的64%至81%。基于Tapas的FMM的多GPU版本在具有300个GPU的TSUBAME2.5的100个节点上执行时实现了5.15倍的加速比。
Tapas is our new C++ programming framework for hierarchical algorithms such as N-body, on large scale heterogeneous supercomputers. Although N-body and their variants are widely used in scientific applications, their correct implementations are often difficult on such modern machines, as the algorithms are irregular, complex, and involve explicit task parallel programming over distributed nodes. Encapsulating the complexities in a library or a framework has been challenging due to irregular data access over massively distributed memory. Tapas solves this by converting the users clean implicit-style parallel program into an inspector-executor style code on heterogeneous multi-core, multi-node environment solely by the use of C++ template metaprogramming. A prototype implementation of the Fast Multipole Method on Tapas demonstrates a comparable performance and scaling as ExaFMM, the fastest hand-tuned implementation of FMM, as well as efficient usage of hundreds of GPUs. Specifically, the serial performance is 95% of ExaFMM, whereas the distributed-memory strong-scaling evaluation using up to 1500 CPU cores demonstrates 64% to 81% of the ExaFMM performance. The multi-GPU version of the Tapas-based FMM achieves a 5.15x speedup when executed on 100 nodes of TSUBAME2.5 with 300 GPUs.