课题基金 / 基金详情

PPoSS: Planning: Cross-Layer Design for Cost-Effective HPC in the Cloud

PPoSS: Planning: Cross-Layer Design for Cost-Effective HPC in the Cloud
PPoSS:规划:云中经济高效 HPC 的跨层设计
批准号:
2028929
负责人:
Mahmut Kandemir
金额:
$25.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2022-09-30

项目摘要

项目成果

Mahmut Kandemir的其他基金

相似基金

相关文献

中文摘要
翻译
许多具有国家重要性的高性能计算(HPC)应用(例如,核模拟、气候建模、药物发现、流行病学和金融)处理巨大的数据集,并且具有显著的资源需求和严格的性能/精度/功率约束。不断变化的硬件元件(例如,新兴的新计算元件)和系统软件(对操作系统、编译器和运行时系统的持续修复)使得在本地管理的计算平台中托管这样的HPC应用越来越没有吸引力。 一种有前途的替代方法是将这些应用程序托管在云中。 然而,使传统HPC应用程序云就绪并为给定应用程序确定最佳的云服务组合是需要解决的重大挑战。在这个项目中,采取了一个整体的,跨层的方法来解决的问题,安全地安装在云中的高性能计算应用程序的高效率,低成本,和良好的性能。该项目的一个关键区别在于,它结合了编译时和运行时的创新,并为客户端和云提供商做出了贡献。该项目涵盖以下五个互补的重点,所有这些都是由感兴趣的HPC应用程序的复杂性和规模的增加,以及云服务产品和应用程序服务级别目标的复杂性所带来的挑战:(i)在无数云基础设施选项上表征HPC应用程序行为;(ii)对HPC应用程序云化的编译器支持;(iii)新型编程语言支持--对象即服务(OaaS);(iv)工作负载放置和调度支持;以及(v)对异构硬件上的PaaS/SaaS的系统软件支持。 该项目的最终目标是设计系统的方法,将HPC应用程序映射到多云/混合云中的不同类型的服务(跨越IaaS,SaaS,FaaS,OaaS)。 这项研究有助于改善运行HPC应用程序的成本。该项目还可以轻松地将HPC应用程序从一个云迁移到另一个云,并为云架构设计人员提供数据,以便更好地调整其系统以适应当前和未来的HPC工作负载。除了技术贡献外,该项目还涉及各种教育和外联活动。特别是,一个新的研究生课程,云计算的重点是高性能计算应用程序的创建和免费传播。 该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many high-performance computing (HPC) applications of national importance (e.g., nuclear simulations, climate modeling, drug discovery, epidemiology, and finance) process enormous datasets and have significant resource demands and strict performance/accuracy/power constraints. Ever-changing hardware elements (e.g., emerging new compute elements) and systems software (continuous fixes to operating systems, compilers and runtime systems) make hosting such HPC applications in locally-managed compute platforms increasingly less attractive. A promising alternate approach is to host these applications in the cloud. However, making legacy HPC applications cloud-ready and identifying the best blend of cloud services for a given application are significant challenges that need to be addressed. In this project, a holistic, cross-layer approach is taken to address the problem of securely mounting such HPC applications in the cloud with high efficiency, low cost, and good performance. A key distinguishing aspect of this project is that it combines both compile-time and run-time innovations and makes contributions to both client and cloud-provider sides. This project spans the following five complementary thrusts, all of which are made challenging by the increasing complexity and scale of the HPC applications of interest, and by the complexity of cloud service offerings and application service-level objectives: (i) characterizing HPC application behavior on myriad cloud infrastructural options; (ii) compiler support for HPC application cloudization; (iii) novel programming language support -- Object-as-a-Service (OaaS); (iv) workload placement and scheduling support; and (v) systems software support for PaaS/SaaS on heterogeneous hardware. The ultimate goal of this project is to devise systematic methodologies for mapping HPC applications to different types of services (spanning IaaS, SaaS, FaaS, OaaS) in multi/hybrid-cloud. This research facilitates improvements in the costs of running HPC applications. This project also enables easy transitioning of HPC applications from one cloud to another and provides data for cloud architecture designers to tune their systems better for current and future HPC workloads. In addition to its technical contributions, this project involves various educational and outreach activities as well. In particular, a new graduate curriculum for cloud computing focusing on HPC applications is created and freely disseminated. Finally, the code being developed and experimental results collected are documented and open-sourced.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3472883.3486999
发表时间: 2021-11
期刊: Proceedings of the ACM Symposium on Cloud Computing
影响因子: --
作者: [A. F. Baarzi;G. Kesidis]
通讯作者: A. F. Baarzi;G. Kesidis
DOI: 10.1109/ccgrid54584.2022.00021
发表时间: 2022-05
期刊: 2022 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid)
影响因子: --
作者: [Myungjun Son;S. Mohanty;Jashwant Raj Gunasekaran;Aman Jain;M. Kandemir;G. Kesidis;B. Urgaonkar]
通讯作者: Myungjun Son;S. Mohanty;Jashwant Raj Gunasekaran;Aman Jain;M. Kandemir;G. Kesidis;B. Urgaonkar
Collaborative Research: CNS Core: Small: Resource-efficient, Strongly Consistent Replication for the Cloud
SaTC: CORE: Small: Automatic Software Patching against Microarchitectual Attacks
SHF: Small: Characterizing and Optimizing 3D NAND Flash
Frameworks: Re-Engineering Galaxy for Performance, Scalability and Energy Efficiency
海外基金