On-the-fly Vertex Reuse for Massively-Parallel Software Geometry Processing

On-the-fly Vertex Reuse for Massively-Parallel Software Geometry Processing
复制标题

DOI:
10.1145/3233303
复制
发表时间:
2018-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Michael Kenzel;B. Kerbl;W. Tatzgern;E. Ivanchenko;D. Schmalstieg;M. Steinberger
Michael Kenzel;B. Kerbl;W. Tatzgern;E. Ivanchenko;D. Schmalstieg;M. Steinberger
中科院分区:
其他
文献类型:
--
作者:
Michael Kenzel;B. Kerbl;W. Tatzgern;E. Ivanchenko;D. Schmalstieg;M. Steinberger

文献摘要

被引文献

相似文献

由于其灵活性,计算模式作为实现最先进的渲染流水线的许多算法的一种方式正变得越来越有吸引力。在图形应用中经常遇到的一个关键问题是流顶点和几何处理。在典型的三角形网格中,相同的顶点平均被引用六次。为了避免渲染过程中的冗余计算,传统上使用变换后缓存来重复使用顶点处理结果。然而,这样的顶点高速缓存通常不能在软件中有效地实现,并且随着并行度的增加而不能很好地扩展。我们探索了在大规模并行软件几何处理过程中动态重用逐顶点结果的替代策略。在给定被分成批的输入流的情况下,我们分析了排序、散列和线程组内通信在识别和利用本地重用潜力方面的有效性。我们设计并提出了四种适合现代GPU体系结构的顶点重用策略。我们证明,在各种应用中,这些策略不仅实现了顶点处理结果的有效重用,而且与幼稚方法相比,可以将性能提高2-3倍。奇怪的是,我们的实验还表明,我们的基于批处理的方法表现出与当前图形硬件上的OpenGL实现类似的行为。
Due to its flexibility, compute mode is becoming more and more attractive as a way to implement many of the algorithms part of a state-of-the-art rendering pipeline. A key problem commonly encountered in graphics applications is streaming vertex and geometry processing. In a typical triangle mesh, the same vertex is on average referenced six times. To avoid redundant computation during rendering, a post-transform cache is traditionally employed to reuse vertex processing results. However, such a vertex cache can generally not be implemented efficiently in software and does not scale well as parallelism increases. We explore alternative strategies for reusing per-vertex results on-the-fly during massively-parallel software geometry processing. Given an input stream divided into batches, we analyze the effectiveness of sorting, hashing, and intra-thread-group communication for identifying and exploiting local reuse potential. We design and present four vertex reuse strategies tailored to modern GPU architectures. We demonstrate that, in a variety of applications, these strategies not only achieve effective reuse of vertex processing results, but can boost performance by up to 2-3x compared to a naïve approach. Curiously, our experiments also show that our batch-based approaches exhibit behavior similar to the OpenGL implementation on current graphics hardware.