Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding

Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding
复制标题

DOI:
10.1109/hpca56546.2023.10071054
复制
发表时间:
2023-02
期刊:
2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Bingyao Li;Jieming Yin;Anup Holey;Youtao Zhang;Jun Yang;Xulong Tang
Bingyao Li;Jieming Yin;Anup Holey;Youtao Zhang;Jun Yang;Xulong Tang
中科院分区:
其他
文献类型:
--
作者:
Bingyao Li;Jieming Yin;Anup Holey;Youtao Zhang;Jun Yang;Xulong Tang

文献摘要

被引文献

相似文献

多GPU系统已经成为一种流行的平台,以满足不断增长的应用需求。然而,使用多个GPU并不能保证性能的成比例提高。虽然先前的工作已经广泛研究了优化,以减轻非均匀存储器访问(NUMA)的开销,地址转换过程中也发挥了重要作用,在塑造整体执行性能。本文研究了统一虚拟内存(UVM)下多GPU系统中的地址转换过程。我们特别关注页表遍历的效率,并确定三个主要的延迟损失:i)排队可用页表遍历线程,ii)页表遍历缓存未命中的内存访问,以及iii)处理页错误。根据我们的观察,我们提出了trans-FW,它通过利用大量的翻译共享和渴望远程翻译转发来缩短页表遍历。在10个有代表性的多GPU应用程序上的实验结果表明,我们提出的方法平均提高了53.8%的整体性能。
Multi-GPU systems have become a popular platform to meet the ever-growing application demands. However, employing multiple GPUs does not guarantee proportional performance improvements. While prior works have extensively studied the optimizations to mitigate the non-uniform memory accesses (NUMA) overheads, the address translation process also plays an important role in shaping the overall execution performance. In this paper, we investigate the address translation process in multi-GPU systems under unified virtual memory (UVM). We specifically focus on the efficiency of page table walk and identify three major latency penalties: i) queuing for available page table walk threads, ii) memory accesses for page walk cache misses, and iii) handling page faults. Based on our observations, we propose Trans-FW, which short circuits the page table walk by leveraging substantial translation sharing and eager remote translation forwarding. Experimental results on 10 representative multi-GPU applications show that our proposed approach improves the overall performance by 53.8% on average.