Parallel Algorithms for Successive Convolution

Parallel Algorithms for Successive Convolution
复制标题

连续卷积的并行算法

DOI:
10.1007/s10915-020-01359-x
复制
发表时间:
2021
影响因子:
2.5
通讯作者:
Thavappiragasm, Mathialakan
Thavappiragasm, Mathialakan
中科院分区:
数学2区
文献类型:
--
作者:
Christlieb, Andrew J.;Guthrey, Pierson T.;Sands, William A.;Thavappiragasm, Mathialakan

文献摘要

参考文献

被引文献

相似文献

随着现代计算体系结构的发展,并行性不断增加,这使得以前难以解决的问题在各种科学学科中得以解决。尽管取得了这些进步,但多尺度计算问题仍然对现代架构构成了难以置信的挑战,因为它们需要解决在空间和时间上经常以数量级变化的尺度。这种复杂性导致我们考虑使用涉及积分算子的展开式来近似空间导数的偏微分方程(PDEs)的替代离散化(Christlieb et al. in J computer Phys 379:214 - 236,2019; Christlieb et al.)。计算机科学学报,82 (3):1-29,2020;Christlieb等人。[J] .计算机工程学报,2016,33(5):1145 - 1145。这些构造使用积分项内的显式信息,但隐式地处理边界数据,这有助于提高方法的总体速度。该方法对于线性问题是无条件稳定的,对于非线性问题也是无条件稳定的。此外,它是无矩阵的,即不需要对线性系统进行逆求,也不需要对非线性项进行迭代。此外,该方案采用快速求和算法,计算复杂度为,其中为沿坐标方向的网格点数。虽然已经做了很多工作来探索这些方法背后的理论,但它们在大规模计算环境中的实用性在很大程度上是一个未被探索的话题。在这项工作中,我们通过开发适用于分布式内存系统的域分解算法以及共享内存算法来探索这些方法的性能。作为第一步,我们推导了一个人工的Courant-Friedrichs-Lewy条件,该条件强制采用最近邻(N-N)通信模式,并简要讨论了可能的推广。我们还分析了几种通过优化主要循环结构和最大化数据重用来实现并行算法的方法。采用MPI和Kokkos (Edwards和Trott在J Parallel Distrib computer 74:3202 - 3216,2014)的混合设计分别用于算法的分布式和共享内存组件,我们证明了我们的方法是有效的,并且可以保持更新速度/节点/秒。我们通过几个不同的PDE测试问题,包括一个采用自适应时间步进规则的非线性示例,提供了证明我们算法的可扩展性和多功能性的结果。
The development of modern computing architectures with ever-increasing amounts of parallelism has allowed for the solution of previously intractable problems across a variety of scientific disciplines. Despite these advances, multiscale computing problems continue to pose an incredible challenge to modern architectures because they require resolving scales that often vary by orders of magnitude in both space and time. Such complications have led us to consider alternative discretizations for partial differential equations (PDEs) which use expansions involving integral operators to approximate spatial derivatives (Christlieb et al. in J Comput Phys 379:214–236, 2019; Christlieb et al. J Sci Comput 82:52(3):1–29, 2020; Christlieb et al. J Comput Phys 415:1–25, 2020). These constructions use explicit information within the integral terms, but treat boundary data implicitly, which contributes to the overall speed of the method. This approach is provably unconditionally stable for linear problems and stability has been demonstrated experimentally for nonlinear problems. Additionally, it is matrix-free in the sense that it is not necessary to invert linear systems and iteration is not required for nonlinear terms. Moreover, the scheme employs a fast summation algorithm that yields a method with a computational complexity of, whereNis the number of mesh points along a coordinate direction. While much work has been done to explore the theory behind these methods, their practicality in large scale computing environments is a largely unexplored topic. In this work, we explore the performance of these methods by developing a domain decomposition algorithm suitable for distributed memory systems along with shared memory algorithms. As a first pass, we derive an artificial Courant–Friedrichs–Lewy condition that enforces a nearest-neighbor (N-N) communication pattern and briefly discuss possible generalizations. We also analyze several approaches for implementing the parallel algorithms by optimizing predominant loop structures and maximizing data reuse. Using a hybrid design that employs MPI and Kokkos (Edwards and Trott in J Parallel Distrib Comput 74:3202–3216, 2014) for the distributed and shared memory components of the algorithms, respectively, we show that our methods are efficient and can sustain an update rateDOF/node/s. We provide results that demonstrate the scalability and versatility of our algorithms using several different PDE test problems, including a nonlinear example, which employs an adaptive time-stepping rule.
RAJA 可移植层:概述和现状
DOI: 10.2172/1169830
发表时间: 2014
影响因子: 3.1
作者:
R. Hornung;J. Keasler
通讯作者: J. Keasler
DOI: 10.1007/s10915-020-01152-w
发表时间: 2017-07
影响因子: 2.5
作者:
A. Christlieb;Wei Guo;Yan Jiang;Hyoseon Yang
通讯作者: A. Christlieb;Wei Guo;Yan Jiang;Hyoseon Yang
DOI: 10.1016/j.jcp.2006.03.021
发表时间: 2006-11
期刊: J. Comput. Phys.
影响因子: --
作者:
Lexing Ying;G. Biros;D. Zorin
通讯作者: Lexing Ying;G. Biros;D. Zorin