A Massively Parallel and Scalable Multi-GPU Material Point Method
A Massively Parallel and Scalable Multi-GPU Material Point Method
复制标题
一种大规模并行、可扩展的多GPU质点方法
DOI:
10.1145/3386569.3392442
复制
发表时间:
2020-07
影响因子:
6.2
通讯作者:
Jiang Chenfanfu
中科院分区:
文献类型:
--
作者:
Wang Xinlei;Qiu Yuxing;Slattery Stuart R.;Fang Yu;Li Minchen;Zhu Song-Chun;Zhu Yixin;Tang Min;Manocha Dinesh;Jiang Chenfanfu
Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100X per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.
登录
查看更多内容
DOI:
10.1145/2601097.2601176
发表时间:
2014-07
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
作者:
A. Stomakhin;Craig A. Schroeder;Chenfanfu Jiang;Lawrence Chai;J. Teran;Andrew Selle
通讯作者:
A. Stomakhin;Craig A. Schroeder;Chenfanfu Jiang;Lawrence Chai;J. Teran;Andrew Selle
DOI:
10.1145/3272127.3275082
发表时间:
2018-12
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
作者:
Chu Han;Q. Wen;Shengfeng He;Qianshu Zhu;Yinjie Tan;Guoqiang Han;T. Wong
通讯作者:
Chu Han;Q. Wen;Shengfeng He;Qianshu Zhu;Yinjie Tan;Guoqiang Han;T. Wong
DOI:
10.1145/1730804.1730807
发表时间:
2010-02
期刊:
--
影响因子:
--
作者:
Jonathan M. Cohen;S. Tariq;Simon Green
通讯作者:
Jonathan M. Cohen;S. Tariq;Simon Green
DOI:
10.1145/3173551
发表时间:
2018-06
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
作者:
Omid Mashayekhi;Chinmayee Shah;Hang Qu;Andrew Lim;P. Levis
通讯作者:
Omid Mashayekhi;Chinmayee Shah;Hang Qu;Andrew Lim;P. Levis
影响因子:
2.5
作者:
Chinmayee Shah;David Hyde;Hang Qu;P. Levis
通讯作者:
Chinmayee Shah;David Hyde;Hang Qu;P. Levis