Data‐driven execution of fast multipole methods

Data‐driven execution of fast multipole methods
复制标题

DOI:
10.1002/cpe.3132
复制
发表时间:
2012-03
期刊:
Concurrency and Computation: Practice and Experience
影响因子:
--
通讯作者:
H. Ltaief;Rio Yokota
H. Ltaief;Rio Yokota
中科院分区:
其他
文献类型:
--
作者:
H. Ltaief;Rio Yokota

文献摘要

被引文献

相似文献

快速多极方法 (FMM) 的复杂度为 O(N),受计算限制,并且需要很少的同步,这使得它们成为下一代超级计算机上的有利算法。它们最常见的应用是加速 N 体问题,但也可用于求解边界积分方程。当粒子分布不规则且树结构具有自适应性时,负载平衡就成为一个重要问题。负载平衡 FMM 的常见策略是使用上一步的工作负载作为权重来静态重新分配下一步。作者在论文中讨论了另一种基于数据驱动执行的方法,以有效解决这一具有挑战性的负载平衡问题。核心思想包括将 FMM 最耗时的阶段分解为更小的任务。然后,该算法可以表示为有向无环图,其中节点表示任务,边表示任务之间的依赖关系。该算法的执行是通过使用内核运行时环境的排队和运行时异步调度任务来执行的,以不违反数值正确性目的的数据依赖性的方式执行。这种异步调度会导致无序执行。数据驱动的 FMM 执行的性能结果优于之前的策略,并在四插槽四核 Intel Xeon 系统上显示出线性加速。版权所有 © 2013 John Wiley & Sons, Ltd.
Fast multipole methods (FMMs) have O (N) complexity, are compute bound, and require very little synchronization, which makes them a favorable algorithm on next‐generation supercomputers. Their most common application is to accelerate N‐body problems, but they can also be used to solve boundary integral equations. When the particle distribution is irregular and the tree structure is adaptive, load balancing becomes a non‐trivial question. A common strategy for load balancing FMMs is to use the work load from the previous step as weights to statically repartition the next step. The authors discuss in the paper another approach based on data‐driven execution to efficiently tackle this challenging load balancing problem. The core idea consists of breaking the most time‐consuming stages of the FMMs into smaller tasks. The algorithm can then be represented as a directed acyclic graph where nodes represent tasks and edges represent dependencies among them. The execution of the algorithm is performed by asynchronously scheduling the tasks using the queueing and runtime for kernels runtime environment, in a way such that data dependencies are not violated for numerical correctness purposes. This asynchronous scheduling results in an out‐of‐order execution. The performance results of the data‐driven FMM execution outperform the previous strategy and show linear speedup on a quad‐socket quad‐core Intel Xeon system.Copyright © 2013 John Wiley & Sons, Ltd.