Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB Design

Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB Design
复制标题

DOI:
10.1145/3466752.3480083
复制
发表时间:
2021-10
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Bingyao Li;Jieming Yin;Youtao Zhang;Xulong Tang
Bingyao Li;Jieming Yin;Youtao Zhang;Xulong Tang
中科院分区:
其他
文献类型:
--
作者:
Bingyao Li;Jieming Yin;Youtao Zhang;Xulong Tang

文献摘要

相似文献

近年来,不断增长的应用程序复杂性和输入数据集大小已经推动了多GPU系统作为许多应用领域的理想计算平台的普及。虽然采用多个GPU直观地暴露了应用程序加速的大量并行性,但所提供的性能很少与GPU的数量成比例。背后的主要挑战之一是地址翻译效率。许多先前的工作集中在CPU或单GPU执行场景,而在多GPU系统中的地址转换很少受到关注。在本文中,我们进行了全面的调查,在“单应用程序-多GPU”和“多应用程序-多GPU”的执行范例的地址转换效率。基于我们的观察,我们提出了一种新的TLB层次设计,称为最小TLB,适合多GPU系统,有效地提高了TLB的性能,以最小的硬件开销。在9个单应用负载和10个多应用负载上的实验结果表明,所提出的最小TLB平均分别提高了23.5%和16.3%的性能。
In recent years, the ever-growing application complexity and input dataset sizes have driven the popularity of multi-GPU systems as a desirable computing platform for many application domains. While employing multiple GPUs intuitively exposes substantial parallelism for the application acceleration, the delivered performance rarely scales with the number of GPUs. One of the major challenges behind is the address translation efficiency. Many prior works focus on CPUs or single GPU execution scenarios while the address translation in multi-GPU systems receives little attention. In this paper, we conduct a comprehensive investigation of the address translation efficiency in both “single-application-multi-GPU” and “multi-application-multi-GPU” execution paradigms. Based on our observations, we propose a new TLB hierarchy design, called least-TLB, tailored for multi-GPU systems and effectively improves the TLB performance with minimal hardware overheads. Experimental results on 9 single-application workloads and 10 multi-application workloads indicate the proposed least-TLB improves the performances, on average, by 23.5% and 16.3%, respectively.