课题基金 / 基金详情

Collaborative Research: Frameworks: Designing Next-Generation MPI Libraries for Emerging Dense GPU Systems

Collaborative Research: Frameworks: Designing Next-Generation MPI Libraries for Emerging Dense GPU Systems
协作研究:框架:为新兴密集 GPU 系统设计下一代 MPI 库
批准号:
1931354
负责人:
William Barth
金额:
$38.32万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-11-01 至 2023-10-31

项目摘要

项目成果

William Barth的其他基金

相似基金

相关文献

中文摘要
翻译
现代图形处理单元(gpu)和高性能互连提供的极高计算和通信能力导致了高性能计算(HPC)平台的创建,每个节点具有多个gpu和高性能互连。不幸的是,流行的消息传递接口(MPI)编程模型的最先进的生产质量实现没有适当的支持来为这种密集的GPU系统上的应用程序提供最佳性能和可伸缩性。高端计算(HEC)技术的这些发展和相关中间件问题导致了以下广泛的挑战:如何增强现有的生产质量MPI中间件,以利用新兴的网络技术,为新兴的密集GPU系统上的HPC和深度学习(DL)应用提供最佳的扩展和扩展?俄亥俄州立大学(OSU)和俄亥俄超级计算机中心(OSC)的计算机科学家以及德克萨斯高级计算中心(TACC)、圣地亚哥超级计算机中心(SDSC)和加州大学圣地亚哥分校(UCSD)的计算科学家提出了一项协同综合研究计划,以创新的解决方案解决上述广泛的挑战。提议的框架将提供给合作者和更广泛的科学界,以了解提议的创新对下一代HPC和DL框架的影响,以及在各个科学领域的应用。该项目将培养多名研究生和本科生成为高性能计算领域的未来科学家和工程师。拟议的工作将通过对俄勒冈州立大学、SDSC和TACC新数据科学项目关键课程的教学法研究,推动课程的进步。TACC、SDSC和OSC已建立的全国性培训和推广方案将用于向XSEDE用户传播这项研究的结果。我们将在PEARC、SC和其他会议上组织教程和研讨会,与社区分享研究成果和经验。该项目与美国国家战略计算计划(NSCI)保持一致,该计划旨在推进美国在高性能计算领域的领导地位,并与美国政府最近提出的保持人工智能(AI)领导地位的倡议保持一致。提出的创新包括:1)设计高性能和可扩展的点对点和集体通信操作,充分利用多个网络适配器和先进的网络内计算特性,为节点内和节点间的GPU和CPU缓冲区;2)设计新颖的数据类型处理和统一的内存管理,提高应用性能;3)设计cuda感知I/O子系统,加速HPC和DL应用的MPI I/O和检查点重启;4)设计对容器化环境的支持,以便更好地在现代云环境中轻松部署所提出的解决方案;5)进行集成开发和评估,以确保提出的设计与驱动应用的适当集成。提出的设计将集成到广泛使用的MVAPICH2库中并提供。项目团队成员将与内部和外部合作者密切合作,以促进已发布软件的广泛部署和采用。提出的解决方案旨在实现在新兴的密集GPU平台上的驱动科学领域(分子动力学、晶格QCD、地震学、图像分类和融合研究)的放大和扩展。提议的开发工作的变革性影响是实现HPC和DL框架和应用程序的可扩展性,性能和可移植性,以利用新兴的密集GPU平台,从而导致科学和工程方面的重大进步。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The extremely high compute and communication capabilities offered by modern Graphics Processing Units (GPUs) and high-performance interconnects have led to the creation of High-Performance Computing (HPC) platforms with multiple GPUs and high-performance interconnects per node. Unfortunately, state-of-the-art production quality implementations of the popular Message Passing Interface (MPI) programming model do not have the appropriate support to deliver the best performance and scalability for applications on such dense GPU systems. These developments in High-End Computing (HEC) technologies and associated middleware issues lead to the following broad challenge: How can existing production quality MPI middleware be enhanced to take advantage of emerging networking technologies to deliver the best possible scale-up and scale-out for HPC and Deep Learning (DL) applications on emerging dense GPU systems? A synergistic and comprehensive research plan, involving computer scientists from The Ohio State University (OSU) and Ohio Supercomputer Center (OSC) and computational scientists from the Texas Advanced Computing Center (TACC), and San Diego Supercomputer Center (SDSC) and University of California San Diego (UCSD), is proposed to address the above broad challenges with innovative solutions. The proposed framework will be made available to collaborators and the broader scientific community to understand the impact of the proposed innovations on next-generation HPC and DL frameworks and applications in various science domains. Multiple graduate and undergraduate students will be trained under this project as future scientists and engineers in HPC. The proposed work will enable curriculum advancements via research in pedagogy for key courses in the new Data Science programs at OSU, SDSC and TACC. The established national-scale training and outreach programs at TACC, SDSC and OSC will be used to disseminate the results of this research to XSEDE users. Tutorials and workshops will be organized at PEARC, SC and other conferences to share the research results and experience with the community. The project is aligned with the National Strategic Computing Initiative (NSCI) to advance US leadership in HPC and the recent initiative of the US Government to maintain leadership in Artificial Intelligence (AI.)The proposed innovations include: 1) Designing high-performance and scalable point-to-point, and collective communication operations that fully utilize multiple network adapters and advanced in-network computing features for GPU and CPU buffers within and across nodes; 2) Designing novel datatype processing and unified memory management to improve application performance; 3) Designing CUDA-aware I/O subsystem to accelerate MPI I/O and checkpoint-restart for HPC and DL applications; 4) Designing support for containerized environments to better enable easy deployment of proposed solutions on modern cloud environments; and 5) Carry out integrated development and evaluation to ensure proper integration of proposed designs with the driving applications. The proposed designs will be integrated into the widely-used MVAPICH2 library and made available. The project team members will work closely with internal and external collaborators to facilitate wide deployment and adoption of released software. The proposed solutions will be targeted to enable scale-up and scale-out of the driving science domains (molecular dynamics, lattice QCD, seismology, image classification, and fusion research) on emerging dense GPU platforms. The transformative impact of the proposed development effort is to achieve scalability, performance, and portability out of HPC and DL frameworks and applications to take advantage of emerging dense GPU platforms and hence, leading to significant advancements in science and engineering.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/tpds.2022.3161187
发表时间: 2022
期刊: IEEE Transactions on Parallel and Distributed Systems
影响因子: 5.3
作者: [J. G. Pauloski;Lei Huang;Weijia Xu;K. Chard;I. Foster;Zhao Zhang]
通讯作者: J. G. Pauloski;Lei Huang;Weijia Xu;K. Chard;I. Foster;Zhao Zhang
DOI: 10.1145/3458817.3476152
发表时间: 2021-07
期刊: SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者: [J. G. Pauloski;Qi Huang;Lei Huang;S. Venkataraman;K. Chard;Ian T. Foster;Zhao Zhang]
通讯作者: J. G. Pauloski;Qi Huang;Lei Huang;S. Venkataraman;K. Chard;Ian T. Foster;Zhao Zhang
SHF: Large: Collaborative Research: Next Generation Communication Mechanisms exploiting Heterogeneity, Hierarchy and Concurrency for Emerging HPC Systems
  • 批准号:
    1565431
  • 项目类别:
    Standard Grant
  • 资助金额:
    $42.25万
  • 财政年份:
    2016
  • 负责人:
    William Barth
  • 依托单位:
Collaborative Research: Integrated HPC Systems Usage and Performance of Resources Monitoring and Modeling (SUPReMM)
  • 批准号:
    1203604
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.79万
  • 财政年份:
    2012
  • 负责人:
    William Barth
  • 依托单位:
Collaborative Research: SI2-SSI: A Comprehensive Performance Tuning Framework for the MPI Stack
  • 批准号:
    1148424
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2012
  • 负责人:
    William Barth
  • 依托单位:
SHF:Large:Collaborative Research:Unified Runtime for Supporting Hybrid Programming Models on Heterogeneous Architecture
  • 批准号:
    1213057
  • 项目类别:
    Standard Grant
  • 资助金额:
    $37.19万
  • 财政年份:
    2012
  • 负责人:
    William Barth
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)