How to obtain efficient GPU kernels: An illustration using FMM & FGT algorithms

How to obtain efficient GPU kernels: An illustration using FMM & FGT algorithms
复制标题

DOI:
10.1016/j.cpc.2011.05.002
复制
发表时间:
2010-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Felipe A. Cruz;S. Layton;L. Barba
Felipe A. Cruz;S. Layton;L. Barba
中科院分区:
其他
文献类型:
--
作者:
Felipe A. Cruz;S. Layton;L. Barba

文献摘要

被引文献

相似文献

图形处理器上的计算可能是几十年来计算科学中最重要的发展之一。Beowulf集群将开源软件与商用硬件相结合,真正实现了高性能计算的民主化,自从Beowulf集群问世以来,这个社区从未如此电气化过。就像那时,机遇伴随着挑战而来。科学算法的形成需要重新思考核心方法,以利用新体系结构提供的性能。在这里,我们解决了快速求和算法(快速多极子方法和快速高斯变换),并应用算法重新设计,以达到在GPU上的性能。所获得的性能改进的进展说明了为GPU的大规模并行体系结构制定算法的实践。最终结果是在一张C1060卡上运行超过500GOP/S的内核,从而接近实用峰值。
Computing on graphics processors is maybe one of the most important developments in computational science to happen in decades. Not since the arrival of the Beowulf cluster, which combined open source software with commodity hardware to truly democratize high-performance computing, has the community been so electrified. Like then, the opportunity comes with challenges. The formulation of scientific algorithms to take advantage of the performance offered by the new architecture requires rethinking core methods. Here, we have tackled fast summation algorithms (fast multipole method and fast Gauss transform), and applied algorithmic redesign for attaining performance on gpus. The progression of performance improvements attained illustrates the exercise of formulating algorithms for the massively parallel architecture of the gpu. The end result has beengpukernels that run at over 500 Gop/s on onenvidiatesla C1060 card, thereby reaching close to practical peak.