Data structure design for GPU based heterogeneous systems

Data structure design for GPU based heterogeneous systems
复制标题

基于GPU的异构系统数据结构设计

DOI:
--
复制
发表时间:
2009
期刊:
International Symposium on High Performance Computing Systems and Applications
影响因子:
--
通讯作者:
Jens Breitbart
Jens Breitbart
中科院分区:
--
文献类型:
--
作者:
Jens Breitbart

文献摘要

被引文献

相似文献

本文报告了我们在为具有多个CPU核心和可编程图形卡的系统进行数据结构设计方面的经验。我们将我们的数据结构集成到类似游戏的应用程序OpenSteerDemo中,并在两个PC系统上对我们的数据结构进行了比较。一个系统具有相对较快的单核CPU和较慢的GPU,而另一个系统使用高端GPU和较慢的多核CPU。我们设计了两种基于网格的数据结构,以有效解决k - 近邻问题。静态网格使用大小均匀的网格单元,而动态网格不依赖给定的网格单元,而是在运行时创建它们。静态网格旨在快速创建数据结构,而动态网格旨在提供高GPU模拟性能。高性能是以更复杂的构建算法为代价,通过利用GPU内存系统实现的。我们的实验表明,在CPU较慢的情况下,创建动态网格的算法成为瓶颈,与静态网格相比,总体性能无法提高。当使用较快的CPU和较慢的GPU进行模拟时也是如此,尽管平衡点不同。我们尝试在GPU上创建数据结构,但静态网格的性能不可行。由于缺乏对递归函数的支持,动态网格无法在GPU上创建。我们提供了一种使用多个CPU核心的动态网格创建算法。由于并行化开销,该算法比其顺序执行的对应算法慢。
This paper reports on our experience with data structure design for systems having both multiple CPU cores and a programmable graphics card. We integrate our data structures into the game-like application OpenSteerDemo and compare our data structures on two pc-systems. One System has a relative fast single core CPU and slower GPU, whereas the other one uses a high-end GPU with a slower multi core CPU. We design two grid based data structures for effectively solving the k-nearest neighbor problem. The static grid uses grid cells of uniform size, whereas the dynamic grid does not rely on given grid cells, but creates them at runtime. The static grid is designed for fast data structure creation, whereas the dynamic grid is designed to provide high GPU simulation performance. The high performance is achieved by taking advantage of the GPU memory system at the cost of a more complex construction algorithm. Our experiments show that with a slower CPU the algorithm for creating the dynamic grid becomes the bottleneck and no overall performance increase is possible compared to the static grid. This also holds true when the simulation is run with a faster CPU and a slower GPU, even though the breakeven point is different. We experimented with data structure creation on the GPU, but the performance of the static grid is not feasible. The dynamic grid cannot be created on the GPU due to the lack of recursive function support. We provide a dynamic grid creation algorithm, which uses multiple CPU cores. This algorithm is slower than its sequential counterpart due to the parallelization overhead.