A PRELIMINARY INVESTIGATION OF EMULATING APPLICATIONS THAT USE PETABYTES OF MEMORY ON PETASCALE MACHINES

A PRELIMINARY INVESTIGATION OF EMULATING APPLICATIONS THAT USE PETABYTES OF MEMORY ON PETASCALE MACHINES
复制标题

对在 PB 级机器上使用 PB 内存的模拟应用程序的初步研究

DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
Chao Mei
Chao Mei
中科院分区:
--
文献类型:
--
作者:
Chao Mei

文献摘要

被引文献

相似文献

随着千万亿次计算时代的到来,一台典型的千万亿次计算机通常安装有数十万个处理器。为了预测在这样的机器上运行的现有应用程序的性能,基于仿真的方法面临着各种前所未有的挑战。其中之一是内存约束所造成的模拟petascale的应用程序,使用petabytes的内存在现有的并行机,通常只有数百TB的内存或less.With试图解决这个挑战,我们提出了一个初步的调查,利用现有的性能预测框架BigSim的上下文中的核心外的概念。我们设计了三个基本的核外方案,并探讨了可以应用于基本方案的优化。理论分析对有核外支持和无核外支持两种情况下的仿真性能进行了建模,指出了对基本核外方案进行优化的重要性,并对其实现提出了建议。通过解决软件堆栈中的多个问题,我们实现了基本的核心外支持的原型,并将其集成到BigSim的仿真组件中。实验结果表明,该原型系统在Charm++和MPI应用程序上都能正常工作,并在MPI内核基准测试中具有合理的仿真性能。
With the advent of petascale computing era, a typical petascale machine is usually installed with hundreds of thousands of processors. To predict the performance of existing applications running on such a machine, a simulation-based approach faces various unprecedented challenges. One of them is the memory constraint posed by emulating the petascale application that uses petabytes of memory on an existing parallel machine that typically only has hundreds of terabytes of memory or less. With an attempt to address this challenge, we present a preliminary investigation of utilizing the out-of-core concept in the context of an existing performance prediction framework BigSim. We have designed three basic out-of-core schemes and explored the optimizations that could be applied to the basic schemes. The additional theoretical analysis models both the performance of the emulation with and without out-of-core support, points out the very importance of the optimizations for the basic out-of-core schemes and offers suggestion to its implementation. By resolving multiple issues across the software stack, we implemented a prototype of the basic out-of-core support and integrated it into the emulation component of BigSim. The experimental evaluation shows the prototype works correctly for both Charm++ and MPI applications, and a reasonable emulation performance regarding a MPI kernel benchmark.