Portable and Vendor-Independent Low-Level Programming and Performance Benchmarking for Graphics Cards and Processors

Portable and Vendor-Independent Low-Level Programming and Performance Benchmarking for Graphics Cards and Processors
复制标题

适用于显卡和处理器的便携式且独立于供应商的低级编程和性能基准测试

DOI:
--
复制
发表时间:
2017
期刊:
2017 IEEE 19th International Conference on High Performance Computing and Communications Workshops (HPCCWS)
影响因子:
--
通讯作者:
V. Lindenstruth
V. Lindenstruth
中科院分区:
--
文献类型:
--
作者:
D. Rohr;V. Lindenstruth

文献摘要

参考文献

被引文献

相似文献

gpu具有显著加快程序速度的潜力,并且是增加计算密集型科学应用的科学范围的一个机会。基于C和其他语言的几个新的编程模型已经发展到可以利用这种并行架构的潜力。但是,使用不同的语言和api开发单独的源代码版本会降低代码的可维护性。它还可能导致略有不同的输出,使结果的验证复杂化。为了比较计算性能,必须考虑处理器的不同性质、不同的价格和不同的硬件速度等级。在本文中,我们总结了一组应用程序适应gpu的经验。我们给出了几个案例,说明我们如何为多个体系结构实现通用代码,以及我们如何克服出现的挑战。提出的应用包括一种重建粒子轨迹的算法,用于欧洲核子研究中心大型强子对撞机ALICE高能级触发器;用于对超级计算机的性能进行排名的Linpack基准,特别是它的矩阵乘法子步骤;冗余数据存储中基于Reed-Solomon的故障擦除编码晶格量子色动力学计算;以及电子显微镜图像评价的应用。
GPUs have the potential to speed up programs significantly and are one opportunity to increase the scientific reach of compute-intense scientific applications. Several new programming models based on C and other languages have evolved to leverage the potential of such parallel architectures. Still, the development of individual source code versions using different languages and APIs deteriorates the maintainability of the code. It can also lead to slightly different outputs complicating the verification of the results. For comparing the compute performance, the different nature of the processors, different pricing, and different speed grades of hardware must be taken into account. In this paper, we summarize our experience from adapting a set of applications to GPUs. We present several cases how we implement generic code for multiple architectures and how we overcome the challenges that occurred. The presented applications encompass an algorithm to reconstruct the trajectories of particles for the ALICE High Level Trigger at the Large Hadron Collider at CERN; the Linpack benchmark used for ranking the performance of supercomputers and in particular its matrix multiplication substep; Reed-Solomon based failure erasure coding for redundant data storage; Lattice Quantum Chromo Dynamics computations; and an application for evaluating electron microscopy images.
Lattice-CSC:为Lattice-QCD优化和构建高效超级计算机并在Green500中获得第一名
DOI: 10.1007/978-3-319-20119-1_14
发表时间: 2015
期刊:
影响因子: --
作者:
D. Rohr;M. Bach;G. Nešković;V. Lindenstruth;C. Pinke;O. Philipsen
通讯作者: O. Philipsen