XUnified: A Framework for Guiding Optimal Use of GPU Unified Memory

XUnified: A Framework for Guiding Optimal Use of GPU Unified Memory
复制标题

XUnified:指导 GPU 统一内存优化使用的框架

DOI:
10.1109/access.2022.3196008
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Liao, Chunhua
Liao, Chunhua
中科院分区:
计算机科学3区
文献类型:
--
作者:
Xu, Hailu;Lin, Pei-Hung;Emani, Murali;Hu, Liting;Liao, Chunhua

文献摘要

参考文献

被引文献

相似文献

统一内存是系统中的任何处理器(GPU或CPU)都可以访问的单个内存地址空间。NVIDIA的统一内存在物理上分离的CPU和GPU内存之上创建了一个托管内存池。NVIDIA的统一内存可按需自动迁移页面级数据,因此程序员可以在不同的机器上快速开发CUDA代码。然而,程序员很难决定何时以及如何有效地使用NVIDIA的统一存储器,因为(1)用户通常不知道应该将哪个统一存储器提示(例如,ReadMostly、PferredLocation、AccessedBy)用于应用程序中的数据对象,以及(2)对于具有不同数据对象或输入的各种应用程序进行手动存储器管理(即,手动代码修改)是繁琐且容易出错的。我们提出了一种将离线训练和在线适应相结合的建议控制器XUnified,以指导在运行时为各种应用程序优化使用统一内存和离散内存。离线阶段使用分析器生成的度量来训练机器学习模型,该模型用于预测最佳内存建议选择,然后在运行时将该建议应用于应用程序。我们使用一组不同的计算基准对XUnifiedon NVIDIA Volta GPU进行了评估。实验结果表明,该算法在正确识别最优记忆建议选择方面达到了94.0%的预测正确率,最大可减少34.3%的内核执行时间。
Unified Memory is a single memory address space that is accessible by any processor (GPUs or CPUs) in a system. NVIDIA’s unified memory creates a pool of managed memory on top of physically separated CPU and GPU memories. NVIDIA’s unified memory automatically migrates page-level data on-demand, so programmers can quickly develop CUDA codes on heterogeneous machines. However, it is extremely difficult for programmers to decide when and how to efficiently use NVIDIA’s unified memory because (1) users are usually unaware of which unified memory hint (e.g.,ReadMostly,PreferredLocation,AccessedBy) should be used for a data object in the application, and (2) it is tedious and error-prone to do manual memory management (i.e., manual code modifications) for various applications with difference data objects or inputs. We presentXUnified, anadvice controllerwhich combines the offline training with the online adaptation to guide the optimal use of unified memory and discrete memory for various applications at runtime. The offline phase uses profiler-generated metrics to train a machine learning model, which is used to predict optimal memory advice choice and it then applies this advice to applications at runtime. We evaluateXUnifiedon NVIDIA Volta GPUs with a set of heterogeneous computing benchmarks. Results show that it achieves 94.0% prediction accuracy in correctly identifying the optimal memory advice choice with a maximal 34.3% reduction in kernel execution time.
OpenARC:用于基于指令的加速器编程研究的可扩展 OpenACC 编译器框架
DOI: --
发表时间: 2014
期刊: 2014 First Workshop on Accelerator Programming using Directives
影响因子: --
作者:
Seyong Lee;J. Vetter
通讯作者: J. Vetter
DOI: --
发表时间: 2014
期刊: Journal of management science
影响因子: --
作者:
อนิรุธ สืบสิงห์
通讯作者: อนิรุธ สืบสิงห์
OpenMP GPU 卸载的统一内存基准测试和评估
DOI: 10.1145/3148173.3148184
发表时间: 2017
期刊: Proceedings of the Fourth Workshop on the LLVM Compiler Infrastructure in HPC
影响因子: --
作者:
Alok Mishra;Lingda Li;Martin Kong;H. Finkel;B. Chapman
通讯作者: B. Chapman
ALTIS:现代化 GPGPU 基准测试
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
Bodun Hu;Christopher J. Rossbach
通讯作者: Christopher J. Rossbach