Collaborative Research: Frameworks: Designing Next-Generation MPI Libraries for Emerging Dense GPU Systems
Collaborative Research: Frameworks: Designing Next-Generation MPI Libraries for Emerging Dense GPU Systems
批准号:
1931537
负责人:
Dhabaleswar Panda
金额:
$136.05万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-11-01 至 2023-10-31
中文摘要
现代图形处理单元(GPU)和高性能互连提供的极高计算和通信能力已导致创建每个节点具有多个GPU和高性能互连的高性能计算(HPC)平台。不幸的是,流行的消息传递接口(MPI)编程模型的最先进的生产质量实现没有适当的支持,以提供最佳的性能和可扩展性的应用程序在这种密集的GPU系统。高端计算(HEC)技术和相关中间件问题的这些发展导致了以下广泛的挑战:如何增强现有的生产质量MPI中间件,以利用新兴的网络技术,为新兴的密集GPU系统上的HPC和深度学习(DL)应用提供最佳的纵向扩展和横向扩展?一个协同和全面的研究计划,涉及计算机科学家从俄亥俄州州立大学(OSU)和俄亥俄州超级计算机中心(OSC)和计算科学家从得克萨斯州高级计算中心(TACC),圣地亚哥超级计算机中心(SDSC)和加州圣地亚哥大学(UCSD),提出了解决上述广泛的挑战与创新的解决方案。拟议的框架将提供给合作者和更广泛的科学界,以了解拟议的创新对下一代HPC和DL框架以及各个科学领域应用的影响。 多个研究生和本科生将在这个项目下培养未来的科学家和工程师在HPC。拟议的工作将通过OSU,SDSC和TACC新数据科学计划中的关键课程的教学法研究实现课程进步。在TACC、SDSC和OSC建立的全国范围的培训和推广计划将用于向XSEDE用户传播这项研究的结果。在PEARC、SC和其他会议上将组织讲座和研讨会,与社区分享研究成果和经验。该项目与国家战略计算计划(NSCI)保持一致,以推动美国在HPC领域的领导地位,以及美国政府最近提出的保持人工智能(AI)领导地位的倡议。拟议的创新包括:1)设计高性能和可扩展的点对点和集体通信操作,充分利用多个网络适配器和先进的网络计算功能,用于节点内和节点间的GPU和CPU缓冲区; 2)设计新颖的数据类型处理和统一的内存管理,以提高应用性能; 3)设计CUDA感知的I/O子系统,加速HPC和DL应用的MPI I/O和检查点重启; 4)设计对容器化环境的支持,以更好地在现代云环境中轻松部署建议的解决方案;以及5)进行综合开发和评估,以确保建议的设计与驱动应用程序的适当集成。拟议的设计将被集成到广泛使用的MVAPICH 2库中并提供。项目小组成员将与内部和外部合作者密切合作,以促进广泛部署和采用已发布的软件。提出的解决方案将有针对性地在新兴的密集GPU平台上实现驱动科学领域(分子动力学、晶格QCD、地震学、图像分类和融合研究)的纵向扩展和横向扩展。该项目的发展成果将带来变革性的影响,即实现HPC和DL框架及应用程序的可扩展性、性能和可移植性,从而充分利用新兴的高密度GPU平台,从而推动科学和工程领域的重大进步。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响评审标准进行评估,被认为值得支持。
英文摘要
The extremely high compute and communication capabilities offered by modern Graphics Processing Units (GPUs) and high-performance interconnects have led to the creation of High-Performance Computing (HPC) platforms with multiple GPUs and high-performance interconnects per node. Unfortunately, state-of-the-art production quality implementations of the popular Message Passing Interface (MPI) programming model do not have the appropriate support to deliver the best performance and scalability for applications on such dense GPU systems. These developments in High-End Computing (HEC) technologies and associated middleware issues lead to the following broad challenge: How can existing production quality MPI middleware be enhanced to take advantage of emerging networking technologies to deliver the best possible scale-up and scale-out for HPC and Deep Learning (DL) applications on emerging dense GPU systems? A synergistic and comprehensive research plan, involving computer scientists from The Ohio State University (OSU) and Ohio Supercomputer Center (OSC) and computational scientists from the Texas Advanced Computing Center (TACC), and San Diego Supercomputer Center (SDSC) and University of California San Diego (UCSD), is proposed to address the above broad challenges with innovative solutions. The proposed framework will be made available to collaborators and the broader scientific community to understand the impact of the proposed innovations on next-generation HPC and DL frameworks and applications in various science domains. Multiple graduate and undergraduate students will be trained under this project as future scientists and engineers in HPC. The proposed work will enable curriculum advancements via research in pedagogy for key courses in the new Data Science programs at OSU, SDSC and TACC. The established national-scale training and outreach programs at TACC, SDSC and OSC will be used to disseminate the results of this research to XSEDE users. Tutorials and workshops will be organized at PEARC, SC and other conferences to share the research results and experience with the community. The project is aligned with the National Strategic Computing Initiative (NSCI) to advance US leadership in HPC and the recent initiative of the US Government to maintain leadership in Artificial Intelligence (AI.)The proposed innovations include: 1) Designing high-performance and scalable point-to-point, and collective communication operations that fully utilize multiple network adapters and advanced in-network computing features for GPU and CPU buffers within and across nodes; 2) Designing novel datatype processing and unified memory management to improve application performance; 3) Designing CUDA-aware I/O subsystem to accelerate MPI I/O and checkpoint-restart for HPC and DL applications; 4) Designing support for containerized environments to better enable easy deployment of proposed solutions on modern cloud environments; and 5) Carry out integrated development and evaluation to ensure proper integration of proposed designs with the driving applications. The proposed designs will be integrated into the widely-used MVAPICH2 library and made available. The project team members will work closely with internal and external collaborators to facilitate wide deployment and adoption of released software. The proposed solutions will be targeted to enable scale-up and scale-out of the driving science domains (molecular dynamics, lattice QCD, seismology, image classification, and fusion research) on emerging dense GPU platforms. The transformative impact of the proposed development effort is to achieve scalability, performance, and portability out of HPC and DL frameworks and applications to take advantage of emerging dense GPU platforms and hence, leading to significant advancements in science and engineering.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(36)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Network-Assisted Noncontiguous Transfers for GPU-Aware MPI Libraries
GPU 感知 MPI 库的网络辅助非连续传输
DOI:
10.1109/mm.2023.3241133
发表时间:
2023
期刊:
IEEE Micro
影响因子:
3.6
作者:
[Suresh, Kaushik Kandadi, Khorassani, Kawthar Shafie, Chen, Chen Chun, Ramesh, Bharath, Abduljabbar, Mustafa, Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar K.]
通讯作者:
Panda, Dhabaleswar K.
MCR-DL: Mix-and-Match Communication Runtime for Deep Learning
MCR-DL:用于深度学习的混合匹配通信运行时
DOI:
10.1109/ipdps54959.2023.00103
发表时间:
2023
期刊:
37th IEEE International Parallel & Distributed Processing Symposium
影响因子:
--
作者:
[Anthony, Quentin, Awan, Ammar Ahmad, Rasley, Jeff, He, Yuxiong, Shafi, Aamir, Abduljabbar, Mustafa, Subramoni, Hari, Panda, Dhabaleswar]
通讯作者:
Panda, Dhabaleswar
OMB-Py: Python Micro-Benchmarks for Evaluating Performance of MPI Libraries on HPC Systems
OMB-Py:用于评估 HPC 系统上 MPI 库性能的 Python 微基准
DOI:
10.1109/ipdpsw55747.2022.00143
发表时间:
2022
期刊:
23rd Parallel and Distributed Scientific and Engineering Computing Workshop (PDSEC
影响因子:
--
作者:
[Alnaasan, Nawras, Jain, Arpan, Shafi, Aamir, Subramoni, Hari, Panda, Dhabaleswar K]
通讯作者:
Panda, Dhabaleswar K
DOI:
10.1145/3547276.3548524
发表时间:
2022-08
期刊:
Workshop Proceedings of the 51st International Conference on Parallel Processing
影响因子:
--
作者:
[Tu Tran;Benjamin Michalowicz;B. Ramesh;H. Subramoni;A. Shafi;D. Panda]
通讯作者:
Tu Tran;Benjamin Michalowicz;B. Ramesh;H. Subramoni;A. Shafi;D. Panda
DOI:
10.1109/hipc56025.2022.00025
发表时间:
2022-12
期刊:
2022 IEEE 29th International Conference on High Performance Computing, Data, and Analytics (HiPC)
影响因子:
--
作者:
[K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda]
通讯作者:
K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda
共 33 条
CSR: Small: CONCERT: Designing Scalable Communication Runtimes with On-the-fly Compression for HPC and AI Applications on Heterogeneous Architectures
-
批准号:2312927
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2023
-
负责人:Dhabaleswar Panda
-
依托单位:
Travel: Student Travel Support for MVAPICH User Group (MUG) 2023 Conference
-
批准号:2331223
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2023
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: Frameworks: Performance Engineering Scientific Applications with MVAPICH and TAU using Emerging Communication Primitives
-
批准号:2311830
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2023
-
负责人:Dhabaleswar Panda
-
依托单位:
Travel: Student Travel Support for MVAPICH User group (MUG) 2022 Conference
-
批准号:2231825
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2022
-
负责人:Dhabaleswar Panda
-
依托单位:
AI Institute for Intelligent CyberInfrastructure with Computational Learning in the Environment (ICICLE)
-
批准号:2112606
-
项目类别:Cooperative Agreement
-
资助金额:$2000.0万
-
财政年份:2021
-
负责人:Dhabaleswar Panda
-
依托单位:
MRI: RADiCAL: Reconfigurable Major Research Cyberinfrastructure for Advanced Computational Data Analytics and Machine Learning
-
批准号:2018627
-
项目类别:Standard Grant
-
资助金额:$77.0万
-
财政年份:2020
-
负责人:Dhabaleswar Panda
-
依托单位:
OAC Core: Small: Next-Generation Communication and I/O Middleware for HPC and Deep Learning with Smart NICs
-
批准号:2007991
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2020
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User Group (MUG) Meeting
-
批准号:1930003
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2019
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User Group (MUG) Meeting
-
批准号:1839739
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2018
-
负责人:Dhabaleswar Panda
-
依托单位:
SI2-SSI: FAMII: High Performance and Scalable Fabric Analysis, Monitoring and Introspection Infrastructure for HPC and Big Data
-
批准号:1664137
-
项目类别:Standard Grant
-
资助金额:$80.0万
-
财政年份:2017
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User group (MUG) Meeting
-
批准号:1744956
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2017
-
负责人:Dhabaleswar Panda
-
依托单位:
SHF: Large: Collaborative Research: Next Generation Communication Mechanisms exploiting Heterogeneity, Hierarchy and Concurrency for Emerging HPC Systems
-
批准号:1565414
-
项目类别:Standard Grant
-
资助金额:$117.19万
-
财政年份:2016
-
负责人:Dhabaleswar Panda
-
依托单位:
BD Spokes: SPOKE: MIDWEST: Collaborative:Advanced Computational Neuroscience Network (ACNN)
-
批准号:1636846
-
项目类别:Standard Grant
-
资助金额:$16.65万
-
财政年份:2016
-
负责人:Dhabaleswar Panda
-
依托单位:
Student Travel Support for MVAPICH User Group (MUG) Meeting
-
批准号:1644528
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2016
-
负责人:Dhabaleswar Panda
-
依托单位:
SI2-SSI: Collaborative Research: A Software Infrastructure for MPI Performance Engineering: Integrating MVAPICH and TAU via the MPI Tools Interface
-
批准号:1450440
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2015
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: Chameleon: A Large-Scale, Reconfigurable Experimental Environment for Cloud Research
-
批准号:1419123
-
项目类别:Cooperative Agreement
-
资助金额:$60.0万
-
财政年份:2014
-
负责人:Dhabaleswar Panda
-
依托单位:
BIGDATA: F: DKM: Collaborative Research: Scalable Middleware for Managing and Processing Big Data on Next Generation HPC Systems
-
批准号:1447804
-
项目类别:Standard Grant
-
资助金额:$72.02万
-
财政年份:2014
-
负责人:Dhabaleswar Panda
-
依托单位:
CSR: Eager: HPC Virtualization with SR-IOV
-
批准号:1347189
-
项目类别:Standard Grant
-
资助金额:$9.83万
-
财政年份:2013
-
负责人:Dhabaleswar Panda
-
依托单位:
Collaborative Research: SI2-SSI: A Comprehensive Performance Tuning Framework for the MPI Stack
-
批准号:1148371
-
项目类别:Standard Grant
-
资助金额:$125.16万
-
财政年份:2012
-
负责人:Dhabaleswar Panda
-
依托单位:
SHF: Large: Collaborative Research: Unified Runtime for Supporting Hybrid Programming Models on Heterogeneous Architecture
-
批准号:1213084
-
项目类别:Standard Grant
-
资助金额:$104.58万
-
财政年份:2012
-
负责人:Dhabaleswar Panda
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: