ScELA: scalable and extensible launching architecture for clusters

ScELA: scalable and extensible launching architecture for clusters
复制标题

ScELA:可扩展的集群启动架构

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on High Performance Computing
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
J. K. Sridhar;Matthew J. Koop;Jonathan L. Perkins;D. Panda

文献摘要

被引文献

相似文献

随着集群规模达到数万,当前的作业启动机制无法扩展,因为它们受到资源约束和性能瓶颈的限制。作业启动过程包括两个阶段--在处理器上产生进程和在进程间交换信息以进行作业初始化.各种编程模型的实现遵循不同的信息交换阶段的协议。我们提出了一个可扩展的,可扩展的和高性能的超大规模并行计算joblaunch体系结构的设计。我们提出了这种架构的实现,在10240个处理器内核上启动一个简单的Hello World MPI应用程序时,实现了超过700%的加速比,并且与以前的解决方案相比,处理器内核的数量也增加了3倍以上。
As cluster sizes head into tens of thousands, current joblaunchmechanisms do not scale as they are limited by resource constraintsas well as performance bottlenecks. The job launch process includes twophases - spawning of processes on processors and information exchange betweenprocesses for job initialization. Implementations of various programmingmodels follow distinct protocols for the information exchange phase.We present the design of a scalable, extensible and high-performance joblaunch architecture for very large scale parallel computing. We present implementationsof this architecture which achieve a speedup of more than700% in launching a simple Hello World MPI application on 10, 240 processorcores and also scale to more than 3 times the number of processorcores compared to prior solutions.