RII Track-4: NSF: Scalable MPI with Adaptive Compression for GPU-based Computing Systems
RII Track-4: NSF: Scalable MPI with Adaptive Compression for GPU-based Computing Systems
批准号:
2327266
负责人:
Xin Liang
金额:
$28.07万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-02-01 至 2026-01-31
中文摘要
这项研究基础设施改善轨道4 EPSCoR研究人员项目将为肯塔基大学研究基金会的一名助理教授提供奖学金,并为一名研究生提供培训。这项工作将与阿贡国家实验室(ANL)的研究人员合作进行。消息传递接口(MPI)是在高性能计算系统上执行通信和扩展应用程序的事实标准。MPI的性能对各种下游应用至关重要,包括科学模拟、大数据分析和人工智能。然而,由于gpu的最新发展继续超过商品网络的发展,大尺寸数据传输正在成为最先进的MPI库的主要性能瓶颈。这项工作旨在通过集成数据压缩开发一个高性能和可扩展的MPI库来解决这个问题,这对于充分利用当前和下一代计算系统的能力至关重要。该项目的成功将允许加速科学代码的执行和数据分析,减少运行在基于gpu的大规模计算系统上的应用程序的科学见解的时间。这将有助于在广泛的计算机和计算学科中推进科学发现。该项目的成果将向社会公开,以加强更广泛领域的研究和工程网络基础设施。此外,该项目将通过培养研究生,为先进网络基础设施的教育和劳动力发展做出贡献。该项目旨在提供一个高性能和可扩展的压缩辅助MPI库,以解决高端计算系统中GPU加速器日益增长的计算能力与相对有限的网络带宽之间日益扩大的差距。研究和开发工作将在ANL进行,通过与MPI和科学数据压缩方面的领先专家密切合作,基于他们已建立的软件产品。具体而言,首先将开发一个可组合的GPU压缩框架,该框架以按需构建压缩管道为特征,以便在压缩性能和减小消息大小之间提供平衡的权衡。然后利用这个框架来优化MPI中的点对点通信。之后,将针对科学应用中被认为是主要性能瓶颈的两个重要MPI集合进行量身定制的优化研究,并通过理论分析和实证评估相结合的方式进行彻底的误差量化。此外,实现中将考虑性能可移植性,以适应来自不同供应商的不同体系结构。为此,开发的例程将集成到MPI旗舰库MPICH(消息传递接口变色龙)中,并向研究社区公开提供。交付成果的评估将在ANL的领先计算设备上进行,使用两个关键任务的科学分析。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This Research Infrastructure Improvement Track-4 EPSCoR Research Fellows project will provide a fellowship to an Assistant professor and training for a graduate student at the University of Kentucky Research Foundation. This work will be conducted in collaboration with researchers at the Argonne National Laboratory (ANL). Message Passing Interface (MPI) is the de facto standard to perform communication and scale applications on high-performance computing systems. The performance of MPI is crucial to various downstream applications, including scientific simulations, big data analytics, and artificial intelligence. However, as the recent development of GPUs continues to outpace that of commodity networks, large-size data transfer is becoming the major performance bottleneck in state-of-the-art MPI libraries. This work aims to tackle this problem by developing a performant and scalable MPI library through integrated data compression, which is critical to fully exploit the power of current and next-generation computing systems. The success of this project will allow for accelerated executions of scientific code and data analytics, reducing the time to scientific insights for applications running on large-scale GPU-based computing systems. This will help advance scientific discoveries across a wide range of computer and computational disciplines. The deliverables of this project will be made publicly accessible to the community to enhance the research and engineering cyberinfrastructure in broader domains. In addition, this project will contribute to the education and workforce development for advanced cyberinfrastructure through the training of graduate students.The proposed project aims to deliver a high-performance and scalable compression-assisted MPI library to address the growing gap between the increasing computing power of GPU accelerators and relatively limited network bandwidth in high-end computing systems. The research and development work will be conducted at ANL through close collaborations with the leading experts in MPI and scientific data compression based on their established software products. Specifically, a composable GPU compression framework that features on-demand construction of compression pipeline will be developed first, in order to provide balanced trade-off between compression performance and message size reduction. This framework will then be leveraged to optimize the point-to-point communication in MPI. After that, tailored optimizations will be investigated for two important MPI collectives that are considered the major performance bottlenecks in scientific applications, and thorough error quantization will be performed through a combination of theoretical analysis and empirical evaluation. In addition, performance portability will be considered in the implementation to accommodate for the diverse architectures from different vendors. To this end, the developed routines will be integrated into the flagship MPI library MPICH (Message-Passing Interface Chameleon) and made publicly available to the research community. The evaluation of the deliverables will be performed on the leading computing facilities at ANL using two mission-critical scientific analyses.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: OAC Core: Topology-Aware Data Compression for Scientific Analysis and Visualization
-
批准号:2313122
-
项目类别:Standard Grant
-
资助金额:$20.18万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
Collaborative Research: Elements: ProDM: Developing A Unified Progressive Data Management Library for Exascale Computational Science
-
批准号:2311756
-
项目类别:Standard Grant
-
资助金额:$24.0万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
CRII: OAC: Enabling Quantities-of-Interest Error Control for Trust-Driven Lossy Compression
-
批准号:2330367
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
Collaborative Research: CyberTraining: Pilot: Research Workforce Development for Deep Learning Systems in Advanced GPU Cyberinfrastructure
-
批准号:2330364
-
项目类别:Standard Grant
-
资助金额:$9.87万
-
财政年份:2023
-
负责人:Xin Liang
-
依托单位:
Collaborative Research: CyberTraining: Pilot: Research Workforce Development for Deep Learning Systems in Advanced GPU Cyberinfrastructure
-
批准号:2230098
-
项目类别:Standard Grant
-
资助金额:$9.87万
-
财政年份:2022
-
负责人:Xin Liang
-
依托单位:
CRII: OAC: Enabling Quantities-of-Interest Error Control for Trust-Driven Lossy Compression
-
批准号:2153451
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2022
-
负责人:Xin Liang
-
依托单位:
海外基金