Programming GPGPU Graph Applications with Linear Algebra Building Blocks

Programming GPGPU Graph Applications with Linear Algebra Building Blocks
复制标题

使用线性代数构建模块对 GPGPU 图形应用程序进行编程

DOI:
--
复制
发表时间:
2016
影响因子:
1.5
通讯作者:
S. Reinhardt
S. Reinhardt
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shuai Che;Bradford M. Beckmann;S. Reinhardt

文献摘要

被引文献

相似文献

图应用程序在科学和企业计算中很常见。最近的研究使用图形处理单元(gpu)来加速图形工作负载。这些应用程序倾向于呈现对SIMD执行具有挑战性的特征。为了实现高性能,之前的工作研究了单个图问题,并设计了特定于设备的算法和优化来实现高性能。然而,程序员必须花费大量的手工工作,打包数据和计算以使这种解决方案对gpu友好。对于普通的程序员来说,这通常太复杂了,并且最终的实现可能不具有可移植性和跨平台性能。为了解决这些问题,我们提出并实现了一个带有应用程序示例的软件构建块库,BelRed,它允许程序员轻松构建图形应用程序。BelRed目前建立在OpenCL™框架之上,并针对gpu进行了优化。它由图处理所必需的基本线性代数构建块组成。开发人员可以使用一组关键原语对图算法进行编程。本文介绍了API,并给出了如何使用该库解决各种代表性图问题的几个案例研究。我们在AMD GPU上评估应用程序的性能,并研究优化技术以提高性能。我们表明,该框架有助于为各种图形应用程序提供令人满意的GPU加速,并有助于显著减少编程工作量。
Graph applications are common in scientific and enterprise computing. Recent research used graphics processing units (GPUs) to accelerate graph workloads. These applications tend to present characteristics that are challenging for SIMD execution. To achieve high performance, prior work studied individual graph problems, and designed device-specific algorithms and optimizations to achieve high performance. However, programmers have to expend significant manual effort, packing data and computation to make such solutions GPU-friendly. This usually is too complex for regular programmers, and the resultant implementations may not be portable and perform well across platforms. To address these concerns, we propose and implement a library of software building blocks with application examples, BelRed which allows programmers to build graph applications with ease. BelRed currently is built on top of the OpenCL™ framework and optimized for GPUs. It consists of fundamental linear-algebra building blocks necessary for graph processing. Developers can program graph algorithms with a set of key primitives. This paper introduces the API and presents several case studies on how to use the library for a variety of representative graph problems. We evaluate application performance on an AMD GPU and investigate optimization techniques to improve performance. We show that this framework is useful to provide satisfactory GPU acceleration of various graph applications and help reduce programming efforts significantly.