RII Track-4: NSF: Scalable MPI with Adaptive Compression for GPU-based Computing Systems
RII Track-4: NSF: Scalable MPI with Adaptive Compression for GPU-based Computing Systems
批准号:
2327266
负责人:
Xin Liang
金额:
$28.07万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-02-01 至 2026-01-31
中文摘要
这个研究基础设施改善轨道-4 EPSCoR研究员项目将提供奖学金的助理教授和培训的研究生在肯塔基州研究基金会的大学。这项工作将与阿贡国家实验室(ANL)的研究人员合作进行。消息传递接口(MPI)是在高性能计算系统上执行通信和扩展应用程序的事实标准。MPI的性能对于各种下游应用至关重要,包括科学模拟,大数据分析和人工智能。然而,随着GPU最近的发展继续超过商品网络,大规模数据传输正在成为最先进的MPI库中的主要性能瓶颈。这项工作旨在通过集成数据压缩开发一个高性能和可扩展的MPI库来解决这个问题,这对于充分利用当前和下一代计算系统的能力至关重要。该项目的成功将加速科学代码和数据分析的执行,缩短在基于GPU的大型计算系统上运行的应用程序获得科学见解的时间。这将有助于在广泛的计算机和计算学科中推进科学发现。该项目的成果将向社会公开,以加强更广泛领域的研究和工程网络基础设施。此外,该项目还将通过培训研究生,为先进网络基础设施的教育和劳动力发展做出贡献。拟议项目旨在提供一个高性能和可扩展的压缩辅助MPI库,以解决GPU加速器日益增长的计算能力与高端计算系统相对有限的网络带宽之间日益扩大的差距。研究和开发工作将在ANL通过与MPI和科学数据压缩方面的领先专家的密切合作进行,并基于其既定的软件产品。具体来说,一个可组合的GPU压缩框架,其特点是按需构建压缩管道将首先开发,以提供压缩性能和消息大小减少之间的平衡权衡。然后,将利用此框架来优化MPI中的点对点通信。在此之后,量身定制的优化将被调查的两个重要的MPI集体,被认为是在科学应用中的主要性能瓶颈,并通过理论分析和实证评估相结合,进行彻底的误差量化。此外,在实现中将考虑性能可移植性,以适应不同供应商的不同体系结构。为此,开发的例程将被集成到旗舰MPI库MPICH(消息传递接口变色龙)中,并向研究界公开提供。可交付成果的评估将在ANL领先的计算设施上进行,使用两个关键任务的科学分析。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估。
英文摘要
This Research Infrastructure Improvement Track-4 EPSCoR Research Fellows project will provide a fellowship to an Assistant professor and training for a graduate student at the University of Kentucky Research Foundation. This work will be conducted in collaboration with researchers at the Argonne National Laboratory (ANL). Message Passing Interface (MPI) is the de facto standard to perform communication and scale applications on high-performance computing systems. The performance of MPI is crucial to various downstream applications, including scientific simulations, big data analytics, and artificial intelligence. However, as the recent development of GPUs continues to outpace that of commodity networks, large-size data transfer is becoming the major performance bottleneck in state-of-the-art MPI libraries. This work aims to tackle this problem by developing a performant and scalable MPI library through integrated data compression, which is critical to fully exploit the power of current and next-generation computing systems. The success of this project will allow for accelerated executions of scientific code and data analytics, reducing the time to scientific insights for applications running on large-scale GPU-based computing systems. This will help advance scientific discoveries across a wide range of computer and computational disciplines. The deliverables of this project will be made publicly accessible to the community to enhance the research and engineering cyberinfrastructure in broader domains. In addition, this project will contribute to the education and workforce development for advanced cyberinfrastructure through the training of graduate students.The proposed project aims to deliver a high-performance and scalable compression-assisted MPI library to address the growing gap between the increasing computing power of GPU accelerators and relatively limited network bandwidth in high-end computing systems. The research and development work will be conducted at ANL through close collaborations with the leading experts in MPI and scientific data compression based on their established software products. Specifically, a composable GPU compression framework that features on-demand construction of compression pipeline will be developed first, in order to provide balanced trade-off between compression performance and message size reduction. This framework will then be leveraged to optimize the point-to-point communication in MPI. After that, tailored optimizations will be investigated for two important MPI collectives that are considered the major performance bottlenecks in scientific applications, and thorough error quantization will be performed through a combination of theoretical analysis and empirical evaluation. In addition, performance portability will be considered in the implementation to accommodate for the diverse architectures from different vendors. To this end, the developed routines will be integrated into the flagship MPI library MPICH (Message-Passing Interface Chameleon) and made publicly available to the research community. The evaluation of the deliverables will be performed on the leading computing facilities at ANL using two mission-critical scientific analyses.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: OAC Core: Topology-Aware Data Compression for Scientific Analysis and Visualization
-
批准号:2313122
-
项目类别:Standard Grant
-
资助金额:$20.18万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
Collaborative Research: Elements: ProDM: Developing A Unified Progressive Data Management Library for Exascale Computational Science
-
批准号:2311756
-
项目类别:Standard Grant
-
资助金额:$24.0万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
CRII: OAC: Enabling Quantities-of-Interest Error Control for Trust-Driven Lossy Compression
-
批准号:2330367
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
Collaborative Research: CyberTraining: Pilot: Research Workforce Development for Deep Learning Systems in Advanced GPU Cyberinfrastructure
-
批准号:2330364
-
项目类别:Standard Grant
-
资助金额:$9.87万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
Collaborative Research: CyberTraining: Pilot: Research Workforce Development for Deep Learning Systems in Advanced GPU Cyberinfrastructure
-
批准号:2230098
-
项目类别:Standard Grant
-
资助金额:$9.87万
-
财政年份:2022
-
负责人:Xin Liang
-
依托单位:
CRII: OAC: Enabling Quantities-of-Interest Error Control for Trust-Driven Lossy Compression
-
批准号:2153451
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2022
-
负责人:Xin Liang
-
依托单位:
海外基金