Application-Attuned Memory Management for Containerized HPC Workflows

Application-Attuned Memory Management for Containerized HPC Workflows
复制标题

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
†. MoizArif;†. AvinashMaurya;†. M.MustafaRafique;Dimitrios S. Nikolopoulos;A. R. Butt
†. MoizArif;†. AvinashMaurya;†. M.MustafaRafique;Dimitrios S. Nikolopoulos;A. R. Butt
中科院分区:
其他
文献类型:
--
作者:
†. MoizArif;†. AvinashMaurya;†. M.MustafaRafique;Dimitrios S. Nikolopoulos;A. R. Butt

文献摘要

相似文献

高性能计算(HPC)作业由数据和内存密集型任务组成,通常作为工作流或集成执行,以促进高效和协调的执行。这些工作流传统上在HPC系统上执行,并且基于数据大小、计算复杂性和I/O活动具有独特的内存需求。最近,对这些工作流的容器化执行进行了广泛的探索。HPC作业的容器化工作流执行需要超过节点容量的数tb内存,从而导致过多的数据交换到较慢的存储、作业性能下降和故障。类似地,并置的带宽密集型、延迟敏感或寿命较短的工作流由于争用、内存耗尽和由于次优内存分配而导致的更高的访问延迟而导致性能下降。最近,人们探索了包含持久内存和计算快速链接(CXL)的分层内存系统,以便为内存受限的系统和应用程序提供额外的内存容量和带宽。然而,当前用于分层内存子系统的内存分配和管理技术不足以满足高性能计算系统中并发运行工作流和大规模集成的托管容器化作业的各种需求。本文利用了容器化HPC工作流的分层内存系统,并提出了有效的内存管理策略,包括智能页面放置和退出策略,以提高内存访问性能。我们的页面分配和替换策略包含任务特征,并在工作流之间实现有效的内存共享。我们将我们的策略与流行的HPC调度器SLURM和容器运行时Singularity集成在一起,表明与理想的、现实的和优化的分层执行环境相比,我们的方法提高了分层内存利用率和应用程序性能,并将工作流执行时间分别减少了51%、87%和35%。
—High-Performance Computing (HPC) jobs consist of data and memory-intensive tasks often executed as workflows or ensembles to facilitate efficient and coordinated execution. These workflows are traditionally executed on HPC systems and have unique memory requirements based on the data size, computational complexity, and I/O activity. Recently container-ized execution of these workflows has been extensively explored. Containerized workflow execution of HPC jobs requires several terabytes of memory that exceed node capacity, resulting in excessive data swapping to slower storage, degraded job performance, and failures. Similarly, colocated bandwidth-intensive, latency-sensitive, or short-lived workflows suffer from degraded performance due to contention, memory exhaustion, and higher access latency due to suboptimal memory allocation. Recently, tiered memory systems comprising persistent memory and compute express link (CXL) have been explored to provide additional memory capacity and bandwidth to memory-constrained systems and applications. However, current memory allocation and management techniques for tiered memory subsystems are inadequate to meet the diverse needs of colocated containerized jobs in HPC systems that concurrently run workflows and ensembles at scale. This paper leverages tiered memory systems for containerized HPC workflows and proposes efficient memory management policies including intelligent page placement and eviction policies to improve memory access performance. Our page allocation and replacement policies incorporate task characteristics and enable efficient memory sharing between workflows. We integrate our policies with the popular HPC scheduler, SLURM, and container runtime, Singularity, to show that our approach improves tiered memory utilization and application performance and reduces workflow execution times by up to 51%, 87%, and 35% as compared to the ideal, realistic, and optimized tiered execution environments, respectively.