Enabling Efficient Preemption for SIMT Architectures with Lightweight Context Switching

Enabling Efficient Preemption for SIMT Architectures with Lightweight Context Switching
复制标题

DOI:
10.1109/sc.2016.76
复制
发表时间:
2016-11
期刊:
SC16: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Zhen Lin;L. Nyland;Huiyang Zhou
Zhen Lin;L. Nyland;Huiyang Zhou
中科院分区:
其他
文献类型:
--
作者:
Zhen Lin;L. Nyland;Huiyang Zhou

文献摘要

被引文献

相似文献

上下文切换是一项关键技术,可为CPU提供先发制人的时间和时间培训。但是,对于单个指导多线程(SIMT)处理器(例如高端图形处理单元(GPU)),由于大量线程数量,支持上下文切换是一项挑战在上下文切换期间交换。 SIMT处理器的架构状态包括寄存器,共享内存,SIMT堆栈和障碍状态。最近的作品介绍了Simt处理器上的螺纹块级别的先发制度,以避免上下文切换开销。但是,由于线程块(TB)的执行时间高度取决于内核程序。不能保证先发制的响应时间,并且某些结核病级的抢先技术不能应用于所有内核功能。在本文中,我们提出了三种互补的方法来减少和压缩建筑状态以实现Simt处理器的轻量级上下文。实验表明,我们的方法平均可以将寄存器上下文规模降低91.5%。基于轻巧的上下文切换,我们启用了使用编译器和硬件共同设计对Simt处理器的指令级先发。通过我们提出的方案,与天真的方法相比,预先抢断的潜伏期平均降低了59.7%。
Context switching is a key technique enabling preemption and time-multiplexing for CPUs. However, for single-instruction multiple-thread (SIMT) processors such as high-end graphics processing units (GPUs), it is challenging to support context switching due to the massive number of threads, which leads to a huge amount of architectural states to be swapped during context switching. The architectural state of SIMT processors includes registers, shared memory, SIMT stacks and barrier states. Recent works present thread-block-level preemption on SIMT processors to avoid context switching overhead. However, because the execution time of a thread block (TB) is highly dependent on the kernel program. The response time of preemption cannot be guaranteed and some TB-level preemption techniques cannot be applied to all kernel functions. In this paper, we propose three complementary ways to reduce and compress the architectural states to achieve lightweight context switching on SIMT processors. Experiments show that our approaches can reduce the register context size by 91.5% on average. Based on lightweight context switching, we enable instruction-level preemption on SIMT processors with compiler and hardware co-design. With our proposed schemes, the preemption latency is reduced by 59.7% on average compared to the naive approach.