Machine learning driven system level heterogeneous memory management for high-performance computing
Machine learning driven system level heterogeneous memory management for high-performance computing
批准号:
19K11993
负责人:
GEROFI BALAZS
金额:
$2.75万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2019
资助国家:
日本
项目状态:
已结题
起止时间:
2019-04-01 至 2023-03-31
中文摘要
我们发现,利用机器学习的系统软件级异构内存管理解决方案,特别是基于非监督学习的方法,如强化学习,需要根据跨内存设备的数据布局快速估计执行运行时间,以探索不同的数据放置策略,这使得架构级模拟器对于此目的不切实际。我们提出了一种基于差分跟踪的方法,该方法使用通过基于高频采样的方法(例如,Intel的PEBS)在真实的硬件上使用不同的存储器设备。我们开发了一个运行时估计的基础上,这样的痕迹,提供了一个执行时间估计的数量级快于全系统模拟器。在一些HPC mini应用程序上,我们发现,与真实的硬件上的测量值相比,估计器预测运行时间的平均误差为4.4%。对于深度学习数据洗牌子主题,我们研究了在DL工作者之间划分数据集并在每个训练时期仅执行部分分布式样本交换的可行性。通过在ABCI的2048个GPU和Fugaku的4096个计算节点上进行广泛的实验,我们证明了在实践中,当仔细调整部分分布式交换时,可以保持全局洗牌的验证精度。我们在PyTorch中提供了一个实现,使用户能够控制拟议的数据交换方案。
英文摘要
Results have been achieved in two parallel efforts of the project.We found that system-software-level heterogeneous memory management solutions utilizing machine learning, in particular nonsupervised learning- based methods such as reinforcement learning, require rapid estimation of execution runtime as a function of the data layout across memory devices for exploring different data placement strategies, which renders architecture-level simulators impractical for this purpose. We proposed a differential tracing-based approach using memory access traces obtained by high-frequency sampling-based methods (e.g., Intel's PEBS) on real hardware using of different memory devices. We developed a runtime estimator based on such traces that provides an execution time estimate orders of magnitude faster than full-system simulators. On a number of HPC mini applications we showed that the estimator predicts runtime with an average error of 4.4% compared to measurements on real hardware.For the deep learning data shuffling subtopic, we investigated the viability of partitioning the dataset among DL workers and performing only a partial distributed exchange of samples in each training epoch. Through extensive experiments on up to 2048 GPUs of ABCI and 4096 compute nodes of Fugaku, we demonstrated that in practice validation accuracy of global shuffling can be maintained when carefully tuning the partial distributed exchange. We provided an implementation in PyTorch that enables users to control the proposed data exchange scheme.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/ipdps53621.2022.00109
发表时间:
2022-05
期刊:
2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
作者:
[Thao Nguyen;François Trahay;Jens Domke;Aleksandr Drozd;Emil;Vatai;Jianwei Liao;M. Wahib;]
通讯作者:
Thao Nguyen;François Trahay;Jens Domke;Aleksandr Drozd;Emil;Vatai;Jianwei Liao;M. Wahib;
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Fajardo-Diaz Juan L., Morelos-Gomez Aaron, Cruz-Silva Rodolfo, Matsumoto Akito, Ueno Yutaka, Takeuchi Norihiro, Kitamura Kotaro, Miyakawa Hiroki, Tejima Syogo, Takeuchi Kenji, Tsuzuki Koichi, Endo Morinobu, 田中 紘生,木原 尚,安倍 賢一, Balazs Gerofi]
通讯作者:
Balazs Gerofi
2020 SIAM Conference on Parallel Processing for Scientific Computing
2020 SIAM 科学计算并行处理会议
DOI:
--
发表时间:
2020
期刊:
影响因子:
--
作者:
[宮地英生, 川原慎太郎, Balazs Gerofi]
通讯作者:
Balazs Gerofi
Towards Intelligent Management of Heterogeneous Memory: A Reinforcement Learning Approach
走向异构内存的智能管理:强化学习方法
DOI:
--
发表时间:
2019
期刊:
影响因子:
--
作者:
[宮地英生, 川原慎太郎, 廣渡 祥太,木原 尚,安倍 賢一, Balazs Gerofi]
通讯作者:
Balazs Gerofi
Argonne National Laboratory(米国)
阿贡国家实验室(美国)
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 8 条