Vector-aware register allocation for GPU shader processors

Vector-aware register allocation for GPU shader processors
复制标题

GPU 着色器处理器的向量感知寄存器分配

DOI:
--
复制
发表时间:
2015
期刊:
International Conference on Compilers, Architecture, and Synthesis for Embedded Systems
影响因子:
--
通讯作者:
Szu
Szu
中科院分区:
--
文献类型:
--
作者:
Yi;Szu

文献摘要

被引文献

相似文献

图形处理单元(GPU)现在被广泛用于嵌入式系统中,用于操纵计算机图形,甚至用于通用计算。但是,许多嵌入式系统必须管理高度限制的硬件资源,以实现高性能或能源效率。寄存器的数量是嵌入式GPU设计中常见的限制因素之一。如果未正确设计寄存器分配,则运行较少寄存器的程序可能会遭受高寄存器压力,尤其是在将寄存器分为四个元素的GPU上,并且可以单独访问每个元素,因为为一个寄存器分配寄存器向量型变量不包含所有元素中的值Wastes寄存器空间。在本文中,我们提出了一个矢量意识寄存器分配框架,以改善着着色器体系结构上的注册利用率。该框架涉及两个主要组成部分:(1)基于元素的寄存器分配,根据变量的元素要求分配寄存器,以及(2)注册包装,以增加寄存器的元素以增加连续的自由元素的数量,从而保持更多寄存器中的实时变量。对周期的模拟器的实验结果表明,所提出的框架总共降低了92%的寄存器溢出,并使14个常见的阴暗程序中的91.7%无溢出。这些结果表明,用于存储溢出变量的空间的能源管理的机会,框架将性能提高了溢出变量溢出到内存的一般阴暗处理器的几何平均值,16.3%和29.2%分别具有5-,10和20周期的访问潜伏期。此外,寄存器要求的减少使另外11个具有高寄存器压力的程序可以在轻质的GPU上运行。
Graphics processing units (GPUs) are now widely used in embedded systems for manipulating computer graphics and even for general-purpose computation. However, many embedded systems have to manage highly restricted hardware resources in order to achieve high performance or energy efficiency. The number of registers is one of the common limiting factors in an embedded GPU design. Programs that run with a low number of registers may suffer from high register pressure if register allocation is not properly designed, especially on a GPU in which a register is divided into four elements and each element can be accessed separately, because allocating a register for a vector-type variable that does not contain values in all elements wastes register spaces. In this paper we present a vector-aware register allocation framework to improve register utilization on shader architectures. The framework involves two major components: (1) element-based register allocation that allocates registers based on the element requirement of variables and (2) register packing that rearranges elements of registers in order to increase the number of contiguous free elements, thereby keeping more live variables in registers. Experimental results on a cycle-approximate simulator showed that the proposed framework decreased 92% of register spills in total and made 91.7% of 14 common shader programs spill-free. These results indicate an opportunity for energy management of the space that is used for storing spilled variables, with the framework improving the performance by a geometric mean of 8.3%, 16.3%, and 29.2% for general shader processors in which variables are spilled to memory with 5-, 10-, and 20-cycle access latencies, respectively. Furthermore, the reduction in the register requirement of programs enabled another 11 programs with high register pressure to be runnable on a lightweight GPU.