Pattern-Aware Staging for Hybrid Memory Systems

Pattern-Aware Staging for Hybrid Memory Systems
复制标题

DOI:
10.1007/978-3-030-50743-5_24
复制
发表时间:
2020-05-22
期刊:
High Performance Computing
影响因子:
--
通讯作者:
Schulz M
Schulz M
中科院分区:
其他
文献类型:
--
作者:
Arima E;Schulz M

文献摘要

相似文献

对更高的存储器性能和同时更大的存储器容量的不断增长的需求正将工业引向混合主存储器设计,即,由多种不同的存储技术组成的存储系统。然而,这种趋势自然会导致一个重要的问题:我们如何有效地利用这种混合记忆?我们的论文提出了一种基于软件的方法来解决这一挑战,通过部署模式感知的分期技术。我们的工作是基于以下观察:(a)高带宽的快速存储器优于内存密集型任务的大内存;(B)但这些任务可以运行更长的时间比批量数据复制到/从快速存储器,特别是当访问模式更不规则/稀疏。如果访问是不规则和稀疏的,我们通过应用以下分段技术来利用这些观察结果:(1)将块(几GB的顺序数据)从大内存复制到快速内存;(2)对块执行内存密集型任务;以及(3)将其写回大内存。为了检查在运行时的访问的规律性/稀疏性,可以忽略不计的性能影响,我们开发了一个轻量级的模式检测机制,使用两个不同的布隆过滤器的帮助线程启发的方法。我们的案例研究使用各种科学的代码在一个真实的系统上显示,我们的方法实现了显着的速度相比,只使用大内存或硬件缓存的执行:3或41%的最佳加速,分别。
The ever increasing demand for higher memory performance and—at the same time—larger memory capacity is leading the industry towards hybrid main memory designs, i.e., memory systems that consist of multiple different memory technologies. This trend, however, naturally leads to one important question: how can we efficiently utilize such hybrid memories? Our paper proposes a software-based approach to solve this challenge by deploying a pattern-aware staging technique. Our work is based on the following observations: (a) the high-bandwidth fast memory outperforms the large memory for memory intensive tasks; (b) but those tasks can run for much longer than a bulk data copy to/from the fast memory, especially when the access pattern is more irregular/sparse. We exploit these observations by applying the following staging technique if the accesses are irregular and sparse: (1) copying a chunk (few GB of sequential data) from large to fast memory; (2) performing a memory intensive task on the chunk; and (3) writing it back to the large memory. To check the regularity/sparseness of the accesses at runtime with negligible performance impact, we develop a lightweight pattern detection mechanism using a helper threading inspired approach with two different Bloom filters. Our case study using various scientific codes on a real system shows that our approach achieves significant speed-ups compared to executions with using only the large memory or hardware caching: 3 or 41% speedups in the best, respectively.