A highly available distributed self-scheduler for exascale computing

A highly available distributed self-scheduler for exascale computing
复制标题

DOI:
10.1145/2701126.2701214
复制
发表时间:
2015-01
期刊:
Proceedings of the 9th International Conference on Ubiquitous Information Management and Communication
影响因子:
--
通讯作者:
A. Takefusa;H. Nakada;T. Ikegami;Yoshio Tanaka
A. Takefusa;H. Nakada;T. Ikegami;Yoshio Tanaka
中科院分区:
其他
文献类型:
--
作者:
A. Takefusa;H. Nakada;T. Ikegami;Yoshio Tanaka

文献摘要

相似文献

分层主从模型被认为是一种很有前途的编程范例,为exascale级的高性能计算机。然而,“故障恢复”是艾级计算最重要的问题之一,因为平均故障间隔时间(MTBF)预计是短的。我们提出了一个故障弹性的中间件套件的exascale计算环境。在本文中,我们设计了一个高可用性的分布式自调度作为资源管理系统的建议中间件套件。提出的分布式自调度器由多个进程组成,以实现可扩展性、故障弹性和持久性。我们还开发了一个原型系统的中间件,使用Apache ZooKeeper和Apache Cassandra。使用开发的原型系统的实验表明,建议的分布式自调度器实现所需的故障弹性的应用程序开发使用的中间件,调度器本身也是故障弹性。我们还证实,分布式处理所造成的开销可以减少,调度器可以扩展。
A hierarchical master-worker model is thought to be a promising programming paradigm for exascale-level high performance computers. However, "fault resiliency" is one of the most important issues for exascale computing because the Mean Time Between Failure (MTBF) is expected to be short. We propose a fault resilient middleware suite for exascale computing environments. In this paper, we design a highly available distributed self-scheduler as a resource management system for the proposed middleware suite. The proposed distributed self-scheduler consists of multiple processes in order to achieve scalability, fault resiliency, and persistency. We also develop a prototype system of the middleware, using Apache ZooKeeper and Apache Cassandra. Experiments using the developed prototype system show that the proposed distributed self-scheduler achieves the desired fault resiliency for an application program developed using the middleware, and that the scheduler itself is also fault resilient. We also confirmed that the overheads caused by distributed processing can be reduced, and the scheduler can be scalable.