TASK-BASED FMM FOR MULTICORE ARCHITECTURES

TASK-BASED FMM FOR MULTICORE ARCHITECTURES
复制标题

DOI:
10.1137/130915662
复制
发表时间:
2014-01-01
影响因子:
3.1
通讯作者:
Takahashi, Toru
Takahashi, Toru
中科院分区:
数学2区
文献类型:
--
作者:
Agullo, Emmanuel;Bramas, Berenger;Takahashi, Toru

文献摘要

被引文献

相似文献

快速多极子方法(FMM)是模拟许多物理问题的基本操作。这种方法的高性能设计通常需要针对目标物理和硬件仔细调整算法。在本文中,我们提出了一种新的方法,实现跨架构的高性能。我们的方法包括表示FMM算法作为一个任务流,并采用了国家的最先进的运行时系统,StarPU,处理不同的计算单元上的任务。我们仔细设计了任务流程,数学运算符,它们的实现,和调度方案。在均匀的160核SGI Altix UV 100上,在42.3秒内计算了2亿个粒子上的势和力,并显示出良好的可扩展性。
Fast multipole methods (FMM) are a fundamental operation for the simulation of many physical problems. The high-performance design of such methods usually requires to carefully tune the algorithm for both the targeted physics and the hardware. In this paper, we propose a new approach that achieves high performance across architectures. Our method consists of expressing the FMM algorithm as a task flow and employing a state-of-the-art runtime system, StarPU, to process the tasks on the different computing units. We carefully design the task flow, the mathematical operators, their implementations, and scheduling schemes. Potentials and forces on 200 million particles are computed in 42.3 seconds on a homogeneous 160-core SGI Altix UV 100 and good scalability is shown.