A Massively Parallel and Scalable Multi-GPU Material Point Method

A Massively Parallel and Scalable Multi-GPU Material Point Method
复制标题

一种大规模并行、可扩展的多GPU质点方法

DOI:
10.1145/3386569.3392442
复制
发表时间:
2020-07
影响因子:
6.2
通讯作者:
Jiang Chenfanfu
Jiang Chenfanfu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wang Xinlei;Qiu Yuxing;Slattery Stuart R.;Fang Yu;Li Minchen;Zhu Song-Chun;Zhu Yixin;Tang Min;Manocha Dinesh;Jiang Chenfanfu

文献摘要

参考文献

被引文献

相似文献

利用现代多GPU架构的强大功能,我们提出了一种基于物质点法(MPM)的大规模并行模拟系统,用于模拟经历复杂拓扑变化、自碰撞和大变形的材料的物理行为。我们的系统做出了三个关键贡献。首先,我们引入了一种新的粒子数据结构,它促进了GPU上的合并内存访问模式,并在将粒子数据写入网格时消除了对内存层次结构进行复杂原子操作的需求。其次,我们提出了一种使用新的网格 - 粒子 - 网格(G2P2G)方案的核融合方法,该方法有效地减少了GPU核的启动次数,提高了延迟,并显著减少了存储粒子数据所需的全局内存量。最后,我们引入了优化的算法设计,允许在共享内存环境中使用高效的稀疏网格,使我们能够最好地利用现代多GPU计算平台进行混合拉格朗日 - 欧拉计算模式。我们通过大量的基准测试、评估以及弹塑性、颗粒介质和流体动力学的动态模拟证明了我们方法的有效性。在一个粒子数量从500万到4000万不等的弹性球体碰撞场景中,与一个开源且经过大量优化的基于CPU的MPM代码库[Fang等人,2019]进行比较,我们的GPU MPM在配备英特尔8086K CPU和单个Quadro P6000 GPU的工作站上实现了每时间步超过100倍的加速,这为计算机图形学和计算科学中未来的MPM模拟展现了令人兴奋的可能性。此外,与最先进的GPU MPM方法[Hu等人,2019a]相比,我们不仅在单个GPU上实现了2倍的加速,而且我们的核融合策略和数组 - 结构体 - 数组(AoSoA)数据结构设计也适用于多GPU系统。我们的多GPU MPM在使用4个GPU时表现出近乎完美的弱缩放和强缩放,能够在单个4 - GPU工作站上以每帧不到4分钟的时间在1024³的网格上进行高性能和大规模的模拟(粒子数量接近1亿),在8 - GPU工作站上以每帧不到1分钟的时间模拟1.34亿个粒子。
Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100X per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.
DOI: 10.1145/2601097.2601176
发表时间: 2014-07
期刊: ACM Transactions on Graphics (TOG)
影响因子: --
作者:
A. Stomakhin;Craig A. Schroeder;Chenfanfu Jiang;Lawrence Chai;J. Teran;Andrew Selle
通讯作者: A. Stomakhin;Craig A. Schroeder;Chenfanfu Jiang;Lawrence Chai;J. Teran;Andrew Selle
DOI: 10.1145/3272127.3275082
发表时间: 2018-12
期刊: ACM Transactions on Graphics (TOG)
影响因子: --
作者:
Chu Han;Q. Wen;Shengfeng He;Qianshu Zhu;Yinjie Tan;Guoqiang Han;T. Wong
通讯作者: Chu Han;Q. Wen;Shengfeng He;Qianshu Zhu;Yinjie Tan;Guoqiang Han;T. Wong
DOI: 10.1145/1730804.1730807
发表时间: 2010-02
期刊: --
影响因子: --
作者:
Jonathan M. Cohen;S. Tariq;Simon Green
通讯作者: Jonathan M. Cohen;S. Tariq;Simon Green
DOI: 10.1145/3173551
发表时间: 2018-06
期刊: ACM Transactions on Graphics (TOG)
影响因子: --
作者:
Omid Mashayekhi;Chinmayee Shah;Hang Qu;Andrew Lim;P. Levis
通讯作者: Omid Mashayekhi;Chinmayee Shah;Hang Qu;Andrew Lim;P. Levis
DOI: 10.1111/cgf.13510
发表时间: 2018-09
影响因子: 2.5
作者:
Chinmayee Shah;David Hyde;Hang Qu;P. Levis
通讯作者: Chinmayee Shah;David Hyde;Hang Qu;P. Levis