Optimizing I/O Performance of HPC Applications with Autotuning

Optimizing I/O Performance of HPC Applications with Autotuning
复制标题

DOI:
10.1145/3309205
复制
发表时间:
2019-03-01
影响因子:
1.6
通讯作者:
Snir, Marc
Snir, Marc
中科院分区:
其他
文献类型:
--
作者:
Behzad, Babak;Byna, Surendra;Snir, Marc

文献摘要

被引文献

相似文献

并行输入输出是现代高性能计算(HPC)的重要组成部分。在各种HPC平台上为各种应用程序获得良好的I/O性能是一项重大挑战,部分原因是I/O中间件和硬件之间存在复杂的相互依赖关系。并行文件系统和I/O中间件层都提供了优化参数,理论上可以获得更好的I/O性能。不幸的是,参数的正确组合高度依赖于应用程序、HPC平台、问题大小和并发性。科学应用程序开发人员没有时间或专业知识来承担为每个问题配置识别良好参数的重大负担。他们诉诸于使用系统默认值,这种选择通常会导致I/O性能低下。我们预计这个问题将在exascale级的机器上,这将可能有一个更深的软件栈与分层排列的硬件resources.We提出了一个解决这个问题的自动调优系统,用于优化I/O性能,I/O性能建模,I/O调优,和I/O模式。我们证明了这个框架的价值在几个HPC平台和应用程序的规模。
Parallel Input output is an essential component of modern high-performance computing (HPC). Obtaining good I/O performance for a broad range of applications on diverse HPC platforms is a major challenge, in part, because of complex inter dependencies between I/O middleware and hardware. The parallel file system and I/O middleware layers all offer optimization parameters that can, in theory, result in better I/O performance. Unfortunately, the right combination of parameters is highly dependent on the application, HPC platform, problem size, and concurrency. Scientific application developers do not have the time or expertise to take on the substantial burden of identifying good parameters for each problem configuration. They resort to using system defaults, a choice that frequently results in poor I/O performance. We expect this problem to be compounded on exascale-class machines, which will likely have a deeper software stack with hierarchically arranged hardware resources.We present as a solution to this problem an autotuning system for optimizing I/O performance, I/O performance modeling, I/O tuning, and I/O patterns. We demonstrate the value of this framework across several HPC platforms and applications at scale.