Cloud versus in-house cluster: Evaluating Amazon cluster compute instances for running MPI applications

Cloud versus in-house cluster: Evaluating Amazon cluster compute instances for running MPI applications
复制标题

DOI:
10.1145/2063348.2063363
复制
发表时间:
2011-11
期刊:
2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
Yan Zhai;Mingliang Liu;Jidong Zhai;Xiaosong Ma;Wenguang Chen
Yan Zhai;Mingliang Liu;Jidong Zhai;Xiaosong Ma;Wenguang Chen
中科院分区:
其他
文献类型:
--
作者:
Yan Zhai;Mingliang Liu;Jidong Zhai;Xiaosong Ma;Wenguang Chen

文献摘要

被引文献

相似文献

云服务的出现为构建和使用HPC平台带来了新的可能性。然而,尽管云服务提供了定制的、按需付费的并行计算的灵活性和便利性,但过去三年的多项研究表明,基于云的集群需要显著的性能提升才能成为有竞争力的选择,特别是对于紧密耦合的并行应用程序。在这项工作中,我们研究了在云中运行HPC应用程序的可行性。这项研究在几个方面与现有的调查不同:1)我们对与HPC社区相关的问题进行了全面的研究,包括性能,成本,用户体验和用户活动范围。2)我们将基于Amazon EC2的平台与典型的本地集群和超级计算机选项进行了比较,该平台基于其新推出的面向HPC的虚拟机构建,使用的基准测试和应用程序的规模和问题规模在以前的云HPC研究中是前所未有的。3)我们执行详细的性能和可扩展性分析,以找到最先进的基于云的集群的主要限制因素。4)我们提出了一个案例研究的影响,每个应用程序的并行I/O系统配置唯一启用云服务。我们的研究结果表明,虽然基于EC2的虚拟集群的可扩展性仍然落后于传统的HPC替代品,他们正在迅速获得整体性能和成本效益,使他们成为可行的候选人进行紧密耦合的科学计算。此外,我们详细的基准测试和分析揭示和分析了EC2上的性能和性能稳定性方面的几个问题。
The emergence of cloud services brings new possibilities for constructing and using HPC platforms. However, while cloud services provide the flexibility and convenience of customized, pay-as-you-go parallel computing, multiple previous studies in the past three years have indicated that cloud-based clusters need a significant performance boost to become a competitive choice, especially for tightly coupled parallel applications. In this work, we examine the feasibility of running HPC applications in clouds. This study distinguishes itself from existing investigations in several ways: 1) We carry out a comprehensive examination of issues relevant to the HPC community, including performance, cost, user experience, and range of user activities. 2) We compare an Amazon EC2-based platform built upon its newly available HPC-oriented virtual machines with typical local cluster and supercomputer options, using benchmarks and applications with scale and problem size unprecedented in previous cloud HPC studies. 3) We perform detailed performance and scalability analysis to locate the chief limiting factors of the state-of-the-art cloud based clusters. 4) We present a case study on the impact of per-application parallel I/O system configuration uniquely enabled by cloud services. Our results reveal that though the scalability of EC2-based virtual clusters still lags behind traditional HPC alternatives, they are rapidly gaining in overall performance and cost-effectiveness, making them feasible candidates for performing tightly coupled scientific computing. In addition, our detailed benchmarking and profiling discloses and analyzes several problems regarding the performance and performance stability on EC2.