An Evaluation of DAOS for Simulation and Deep Learning HPC Workloads

An Evaluation of DAOS for Simulation and Deep Learning HPC Workloads
复制标题

DOI:
10.1145/3578353.3589542
复制
发表时间:
2023-05
期刊:
Proceedings of the 3rd Workshop on Challenges and Opportunities of Efficient and Performant Storage Systems
影响因子:
--
通讯作者:
Luke Logan;J. Lofstead;Xian-He Sun;Anthony Kougkas
Luke Logan;J. Lofstead;Xian-He Sun;Anthony Kougkas
中科院分区:
其他
文献类型:
--
作者:
Luke Logan;J. Lofstead;Xian-He Sun;Anthony Kougkas

文献摘要

相似文献

传统上,分布式存储系统依赖于由OS内核提供的接口来与存储硬件交互。然而,许多研究表明,操作系统对每个I/O操作都造成了严重的开销,特别是在高性能存储和网络硬件上(例如,PMEM和200 GBe)。因此,分布式存储堆栈正在被重新设计,以通过利用完全绕过内核的新硬件接口来利用这种现代硬件。然而,这些优化对于真实的硬件上的真实的HPC工作负载的影响尚未得到充分研究。在这项工作中,我们提供了一个全面的评估DAOS:一个国家的最先进的分布式存储系统,重新架构的存储堆栈从头开始为现代硬件。我们将DAOS与传统的存储堆栈进行了比较,并证明了通过利用最佳的硬件接口,在真实的科学应用中可以观察到高达6倍的性能改进。
Traditionally, distributed storage systems have relied upon the interfaces provided by OS kernels to interact with storage hardware. However, much research has shown that OSes impose serious overheads on every I/O operation, especially on high-performance storage and networking hardware (e.g., PMEM and 200GBe). Thus, distributed storage stacks are being re-designed to take advantage of this modern hardware by utilizing new hardware interfaces which bypass the kernel entirely. However, the impact of these optimizations have not been well-studied for real HPC workloads on real hardware. In this work, we provide a comprehensive evaluation of DAOS: a state-of-the-art distributed storage system which re-architects the storage stack from scratch for modern hardware. We compare DAOS against traditional storage stacks and demonstrate that by utilizing optimal interfaces to hardware, performance improvements of up to 6x can be observed in real scientific applications.