Collaborative Research: OAC Core: CropDL - Scheduling and Checkpoint/Restart Support for Deep Learning Applications on HPC Clusters
Collaborative Research: OAC Core: CropDL - Scheduling and Checkpoint/Restart Support for Deep Learning Applications on HPC Clusters
批准号:
2403090
负责人:
Wei Niu
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-10-01 至 2027-09-30
中文摘要
机器学习(ML)和深度学习(DL)(更具体地说,深度神经网络(DNN))工作负载正开始主导高性能计算(HPC)领域。今天,甚至需要大量的计算资源来训练一个最先进的深度学习模型(例如,大型语言模型或LLM)。随着培训大规模DNN模型的需求继续存在,并从私营部门扩展到NSF支持的科学家和工程师(他们更有可能使用共享计算资源),高效的检查点正在成为一项关键需求。检查点不仅有助于处理故障,还为共享的HPC资源提供了更大的调度灵活性,因为非常长的运行时间的作业可以拆分成几个较短的作业。CropDL项目的前提是,高效和自动化的应用级检查点和重启将是促进使用共享HPC集群执行长期运行的ML培训任务的关键,从而大幅增加能够成功培训各种应用的大型ML模型的研究人员数量。该项目还在多个方面促进了教育和多样性,例如,1)引入课程(或课程材料),以引起对计算机系统本科生和研究生教育中与ML相关的工作量的关注;2)将该项目的研究任务与大学的协同研究方案结合起来,以增加妇女和代表性不足的少数群体的参与;3)支持和培训博士生的研究,在与新兴ML工作负载相关的系统和网络基础设施研究方面创造动力,并推广将这些工作负载的属性与现代HPC硬件的复杂性相结合的综合研究。CropDL的总体目标是支持应用程序级别的检查点/深度学习应用程序的重启,以实现更好的弹性、更快的平均完成时间和更高的资源利用率。特别是,DL工作负载的几个特性(与科学计算相比)为检查点创建了截然不同的一组机会和挑战:1)并行执行期间有限的通信模式,这可以实现高效的协调的检查点;2)许多独特的机会来压缩检查点,并且可能采用不协调的检查点;以及3)可延展的执行,其中可能从不同数量的节点重新开始。基于这一观察,本项目的第一个方向是开发检查点设置过程中要训练的DNN模型(S)的性质。这包括在各种并行模型下为DL应用程序设置异步版本化检查点,以及基于内容的数据减少(压缩和稀疏)技术,以减少检查点量。第二个研究方向是如何在设置检查点的同时有效利用现有和即将到来的高性能计算系统的资源。它将来自DL应用程序的任务、数据和I/O需求公式化成DAG表示形式,并开发出调度它们的方法。它还通过新兴的I/O平台为深度学习应用程序支持高效的I/O。最后一个方向是基于DL工作负载的计算图,通过编译系统自动设置检查点。所有这些努力都考虑了DNN的各种并行化方案,即数据、模型和/或流水线并行。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine Learning (ML) and Deep Learning (DL) (more specifically, Deep Neural Network (DNN)) workloads are beginning to dominate the High-Performance Computing (HPC) arena. Today, massive computational resources are required to train even a single state-of-the-art deep learning model (e.g., large language models or LLMs). As the need for training massive DNN models continues and expands from the private sector to NSF-supported scientists and engineers (who are more likely to use shared computing resources), efficient checkpointing is emerging as a critical need. Checkpointing not only helps deal with failures but also provides more scheduling flexibility on shared HPC resources, as a very long-running job can be broken into several shorter ones. The premise of the CropDL project is that efficient and automated application-level checkpoint and restart will be critical to facilitating the use of shared HPC clusters for long-running ML training tasks, drastically increasing the number of researchers that can successfully train large ML models for various applications. This project also contributes to education and diversity in multiple aspects, for example, 1) introducing courses (or course material) to bring attention to ML-related workloads in computer systems undergraduate and graduate education; 2) integrating research tasks from this project with synergistic research programs at universities to increase the participation of women and underrepresented minority groups; and 3) supporting and training PhD students in their research, creating momentum on systems and cyberinfrastructure research related to emerging ML workloads and popularizing integrative research that combines the properties of these workloads with the complexities of modern HPC hardware.The overarching goal of CropDL is to support application-level checkpoints/restarts of deep learning applications for better resiliency, faster average completion time, and higher resource utilization. Particularly, several properties of DL workloads (as compared to scientific computations) create distinct sets of opportunities and challenges for checkpointing: 1) limited communication patterns during parallel execution, which can enable efficient coordinated checkpoints, 2) many unique opportunities for compression of checkpoints, and possibly taking uncoordinated checkpoints, and 3) malleable execution, where restarting from a different number of nodes is possible. Based on this observation, the first direction of this project is to exploit the properties of the DNN model(s) to be trained during checkpointing. This includes asynchronous versioned checkpointing for DL applications under a wide variety of parallelism models as well as content-based data reduction (compression and sparsification) techniques to reduce checkpoint volumes. The second direction of research focuses on using current and upcoming HPC systems' resources efficiently while checkpointing. It formulates tasks, data, and I/O requirements from DL applications into DAG representations and develops methods to schedule them. It also supports efficient I/O for deep learning applications with emerging I/O platforms. The last direction is to automate checkpointing through a compilation system based on the computational graph of DL workloads. All these efforts consider a variety of parallelization schemes for DNNs, i.e., data, model, and/or pipelined parallelism.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Engineering Carboxylic Acid Reductase for the Biosyntheses of Industrial Chemicals
-
批准号:1805528
-
项目类别:Standard Grant
-
资助金额:$33.55万
-
财政年份:2018
-
负责人:Wei Niu
-
依托单位:
SusChEM: Novel 1,2-Propanediol Biosynthesis from Renewable Feedstocks through Enzyme Discovery
-
批准号:1438332
-
项目类别:Standard Grant
-
资助金额:$31.76万
-
财政年份:2014
-
负责人:Wei Niu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: