Utilizing GPU Parallelism to Improve Fast Spherical Harmonic Transforms

Utilizing GPU Parallelism to Improve Fast Spherical Harmonic Transforms
复制标题

DOI:
10.1109/hpec.2018.8547547
复制
发表时间:
2018-09
期刊:
2018 IEEE High Performance extreme Computing Conference (HPEC)
影响因子:
--
通讯作者:
Max Carlson;H. Sundar
Max Carlson;H. Sundar
中科院分区:
其他
文献类型:
--
作者:
Max Carlson;H. Sundar

文献摘要

相似文献

球形谐波构成了生活在球体表面的函数的正交基础,对于求解部分微分方程和数值集成非常有用。将一组函数样本转换为其相应的球形谐波系数的复杂性在很大程度上由相关的Legendre变换的计算主导。这种相关的Legendre变换需要计算$(L+1)$密度矩阵向量产品,其中$ L $是球形谐波扩展的顺序。由于每个矩阵的行数和列的数量都取决于$ L $,因此此步骤本质上是$ o(l^{3})$。在本文中,我们探讨了可用于改善蝴蝶压缩方法的GPU并行性。我们提出了一些初步结果,显示大型问题大小的性能提高,并最终计划释放GPU球形谐波变换的Monarchsht库。
Spherical harmonics form an orthogonal basis for functions that live on the surface of a sphere and are useful for solving partial differential equations and for numerical integration. The complexity of transforming a set of function samples to their corresponding spherical harmonic coefficients is largely dominated by the computation of the associated Legendre transform. This associated Legendre transform requires the computation of $(L+1)$ dense matrix-vector products where $L$ is the order of the spherical harmonic expansion. Since the number of rows and columns of each of these matrices depends on $L$, this step is essentially $O(L^{3})$. In this paper, we explore the GPU parallelism available to improve the butterfly compression approach. We present some preliminary results showing performance increases for large problem sizes and eventually plan to release the MonarchSHT library for GPU spherical harmonic transforms.