SHF: Medium: Collaborative Research: Next-Generation Message Passing for Parallel Programming: Resiliency, Time-to-Solution, Performance-Portability, Scalability, and QoS
SHF: Medium: Collaborative Research: Next-Generation Message Passing for Parallel Programming: Resiliency, Time-to-Solution, Performance-Portability, Scalability, and QoS
批准号:
1562306
负责人:
Sidharth kumar
金额:
$39.79万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-01 至 2022-05-31
中文摘要
基于MPI的并行编程在学术界、政府(国防和非国防用途)以及可扩展机器学习和大数据分析中的新兴用途中的使用频率越来越高。 新兴的超级计算机系统将有更多的故障,MPI需要能够解决这些故障,以适应这些新兴的情况,而不是导致整个应用程序失败。 高性能计算(HPC)的协作,变革性的消息传递研究的关键性能便携式并行编程在新的和即将到来的可扩展系统(与“最佳实践,先,后调试”的战略)正在减少到实践。消息传递接口(MPI-3/4)应用程序编程接口的一个重要子集正在通过具有在并行任务之间同步的弱集体事务的扩展来实现容错。本研究研究的新模型,本地化故障,提供可调的无故障开销,允许多种故障,使分层恢复,是数据并行相关。 正在研究底层网络的故障建模。应用程序开发人员控制这一工作中的粒度和无故障开销。性能和可扩展性的中间件原型的结果主要是通过紧凑的应用程序,涉及到实际和学术兴趣的真实的用例证明。这项工作的影响范围从政府实验室中最大的超级计算机的用户到具有长期运行,时间关键型应用程序的实用集群,以及在故障比过去几年更频繁发生的“敌对”环境中的天基和其他并行处理。 该项目正在生产可用的自由软件,这些软件将在社区中广泛共享,并指导学术界、工业界和政府如何编写更好的并行程序。 该项目还提供了如何更新现有或遗留程序以使用正在减少到实践中的新功能的指导方针。
英文摘要
Parallel programming based on MPI is being used with increased frequency in academia, government (defense and non-defense uses), as well as emerging uses in scalable machine learning and big data analytics. Emerging supercomputer systems will have more faults and MPI needs to be able to workaround such faults to be appropriate to these emerging situations, rather than causing an entire application to fail. Collaborative, transformative message passing research for High Performance Computing (HPC) critical to performance-portable parallel programming in new and forthcoming scalable systems (with a strategy of "best practice-first, standardization-later") is being reduced to practice. A substantial subset of the Message Passing Interface (MPI-3/4) application programmer interface is being made fault tolerant through extensions with weak collective transactions that synchronize between parallel tasks. This research studies the novel model that localizes faults, provides tunable fault-free overhead, allows for multiple kinds of faults, enables hierarchical recovery, and is data-parallel relevant. Fault modeling of underlying networks is being studied. Application developers control the granularity and fault-free overhead in this effort. Performance and scalability results of the middleware prototype are being demonstrated principally through compact applications that relate to real use cases of practical and academic interest. The impact of this work ranges from users of the largest supercomputers in government labs to practical clusters that have long-running, time-critical applications, and to space-based and other parallel processing in "hostile" environments where faults occur more frequently than in past years. The project is producing usable free software that will be widely shared in the community as well as guidance on how better parallel programs can be written in academia, industry, and government. The project also provides guidelines for how to update existing or legacy programs to use the new capabilities that are being reduced to practice.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: SHF: Small: Scalable and Extensible I/O Runtime and Tools for Next Generation Adaptive Data Layouts
-
批准号:2401274
-
项目类别:Standard Grant
-
资助金额:$30.02万
-
财政年份:2023
-
负责人:Sidharth kumar
-
依托单位:
RII Track-4:NSF: Relational Algebra on Heterogeneous Extreme-scale Systems
-
批准号:2132013
-
项目类别:Standard Grant
-
资助金额:$26.48万
-
财政年份:2022
-
负责人:Sidharth kumar
-
依托单位:
Collaborative Research: SHF: Small: Scalable and Extensible I/O Runtime and Tools for Next Generation Adaptive Data Layouts
-
批准号:2221811
-
项目类别:Standard Grant
-
资助金额:$30.02万
-
财政年份:2022
-
负责人:Sidharth kumar
-
依托单位:
海外基金