CAREER : Towards Exascale Performance of Parallel Applications
CAREER : Towards Exascale Performance of Parallel Applications
批准号:
2338077
负责人:
Amanda Bienz
金额:
$55.78万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-02-01 至 2029-01-31
中文摘要
每一代超级计算机都比上一代更强大,目前的系统能够在一秒钟内将数百万个数字相加。这些高性能机器为科学家和工程师提供了解决日益复杂的问题所需的硬件,并通过计算机模拟进行发现。这些程序在组成超级计算机的数千个计算核心上运行,每个核心执行程序的一部分,并根据需要向其他核心传递消息。通常,由于与通信相关联的显著开销,模拟无法有效地使用当前超级计算机提供的计算能力。该项目通过降低广泛使用的并行程序中的通信成本和改进现有应用程序的范围来解决这一挑战,以实现新的科学发现。此外,该项目将支持CS4 ALL课程的振兴,沿着黑客开发,向新墨西哥州大学主校区和分支的不同学生群体介绍计算主题。该项目的目标是最大限度地减少通信成本,提高现有并行应用程序的性能和可扩展性。该项目将为新兴的异构体系结构开发精确的性能模型,以优化非线性求解器,模拟,迭代方法和神经网络中的通信。这些模型将被用来采用几种优化策略,包括图形分区,最大限度地减少性能基于模型的功能,局部感知分区整个广泛使用的代数多重网格(AMG)预处理器,和拓扑感知MPI_Allreduce操作。此外,该项目将探索针对每个节点具有多个GPU的异构系统的专门优化,包括选择最佳通信路径、在通过线程进行通信期间利用所有可用的CPU核心,以及聚合节点间消息以减少注入网络的数据。通过最小化基础数值方法中的通信,该项目的目标是在依赖于这些方法的一系列应用中实现性能和可扩展性的切实改进。2该项目由软件和硬件基金会核心计划和刺激竞争力研究的既定计划(EPSCoR)共同资助该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Each generation of supercomputers is more powerful than the previous, with current systems capable of adding millions of millions of numbers together in a single second. These high-performance machines provide the hardware necessary for scientists and engineers to solve increasingly complex problems and make discoveries through computer simulations. These programs run across the thousands of individual compute cores that make up a supercomputer, with each core executing a portion of the program and communicating messages to other cores as needed. Often, simulations fail to efficiently use the computing power provided by current supercomputers due to significant overheads associated with communication. This project addresses this challenge by reducing communication costs within widely used parallel programs and improving a spectrum of existing applications to allow for novel scientific discoveries. Furthermore, this project will support the revitalization of the CS4ALL course along with hackathon development to introduce computing topics to a diverse student population across the University of New Mexico main and branch campuses.The goals of this project are to minimize communication costs and enhance the performance and scalability of existing parallel applications. This project will develop accurate performance models for emerging heterogeneous architectures to optimize communication within non-linear solvers, simulations, iterative methods, and neural networks. These models will be used to employ several optimization strategies, including graph partitions that minimize performance model-based functions, locality-aware partitioning throughout the widely used algebraic multigrid (AMG) preconditioner, and topology-aware MPI_Allreduce operations. Furthermore, the project will explore specialized optimizations for heterogeneous systems with multiple GPUs per node, including selecting optimal communication paths, utilizing all available CPU cores during communication via threading, and aggregating inter-node messages to reduce data injected into the network. By minimizing communication within foundational numerical methods, this project aims to yield tangible improvements in performance and scalability across a spectrum of applications reliant on these methods.This project is jointly funded by the Software and Hardware Foundations Core Program and the Established Program to Stimulate Competitive Research (EPSCoR).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: EAGER: Real-time Strategies and Synchronized Time Distribution Mechanisms for Enhanced Exascale Performance-Portability and Predictability
-
批准号:2151022
-
项目类别:Standard Grant
-
资助金额:$7.5万
-
财政年份:2022
-
负责人:Amanda Bienz
-
依托单位:
海外基金