To move or not to move?: page migration for irregular applications in over-subscribed GPU memory systems with DynaMap

To move or not to move?: page migration for irregular applications in over-subscribed GPU memory systems with DynaMap
复制标题

DOI:
10.1145/3456727.3463766
复制
发表时间:
2021-06
期刊:
Proceedings of the 14th ACM International Conference on Systems and Storage
影响因子:
--
通讯作者:
Chia-Hao Chang;Adithya Kumar;A. Sivasubramaniam
Chia-Hao Chang;Adithya Kumar;A. Sivasubramaniam
中科院分区:
其他
文献类型:
--
作者:
Chia-Hao Chang;Adithya Kumar;A. Sivasubramaniam

文献摘要

相似文献

本文主要讨论在有限的GPU内存系统上运行大型不规则内存访问应用程序时可能出现的严重页面抖动问题。这种内存超额订阅导致NVIDIA UVM驱动程序中的当前按需(渴望)或基于页面组粒度访问计数器(懒惰)的页面迁移机制的性能非常差。我们对这些执行的详细分析揭示了一个非常新颖的见解:与其在GPU缓存及其内存中重复满足时间和空间局部性的责任,不如前者简单地满足时间方面,后者满足空间方面,从而节省宝贵的内存系统容量。在此基础上,我们构建了一个自适应页面迁移方案,称为DynaMap,该方案(i)使用编译器传递来检测现成的CUDA UVM应用程序的空间利用率跟踪,(ii)动态设置空间利用率阈值以基于内存压力和访问特性确定迁移,以及(iii)增强当前NVIDIA UVM驱动程序以基于阈值动态地迁移页面(从主机存储器到GPU)。使用7个来自公共基准测试套件的不规则应用程序,我们在一个真实的系统上实现了DynaMap,具有不同的超额订阅率,显示速度比最先进的UVM实现高达2.5倍(平均34%)。
This paper focuses on the severe page thrashing problem that can arise when running large irregular memory access applications on limited GPU memory systems. Such memory over-subscription causes very poor performance in the currently on demand (eager) or page-group granularity access-counter based (lazy) page migration mechanisms found in NVIDIA's UVM drivers. Our detailed analysis of these executions reveals a very novel insight: rather than duplicate the responsibility of catering to both temporal and spatial locality in both GPU caches and its memory, it is better for the former to simply cater to the temporal aspect, and the latter to the spatial aspect, thereby saving precious memory system capacities. Based on this, we build an adaptive page migration scheme, called DynaMap, that (i) uses a compiler pass to instrument off-the-shelf CUDA UVM applications for spatial utilization tracking, (ii) dynamically sets a spatial utilization threshold to determine migration based on memory pressure and access characteristics, and (iii) enhances the current NVIDIA UVM driver to dynamically migrate the page (from the host memory to the GPU) based on the threshold. Using 7 irregular applications from public benchmark suites, we implement DynaMap on a real system with different over-subscription ratios to show speedups as much as 2.5X (34% on the average) over state-of-the-art UVM implementations.