课题基金 / 基金详情

CRII: OAC: High-Efficiency Serverless Computing Systems for Deep Learning: A Hybrid CPU/GPU Architecture

CRII: OAC: High-Efficiency Serverless Computing Systems for Deep Learning: A Hybrid CPU/GPU Architecture
CRII:OAC:用于深度学习的高效无服务器计算系统:混合 CPU/GPU 架构
批准号:
2153502
负责人:
Hao Wang
金额:
$17.49万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-05-01 至 2025-04-30

项目摘要

项目成果

Hao Wang的其他基金

相似基金

相关文献

中文摘要
翻译
该奖项全部或部分由《2021年美国救援计划法案》(公法117-2)资助。下一代无服务器云计算为开发人员提供了对服务器管理和管理的简化访问,包括事件驱动的执行、细粒度的资源供应、自动伸缩和按需付费计费。机器学习社区正在利用无服务器云计算的这些优势来简化深度学习(DL)应用程序的开发和部署。然而,现有的无服务器计算平台缺乏对gpu的有效支持,阻碍了深度学习从业者利用无服务器计算进行大规模应用。该项目将开发一个高效的无服务器计算平台,采用混合CPU/GPU架构,以加速DL应用程序的开发和部署。目标是推进深度学习和无服务器计算的前沿方法,这将导致深度学习从业者、深度学习用户和云计算基础设施提供商的重大飞跃,为社会的科学进步做出贡献。研究结果还将通过分布式计算、云计算和深度学习交叉的现实世界系统的令人兴奋的示例和演示来增强本科和研究生教育。该项目将开发一种具有混合CPU/GPU架构的新型无服务器计算平台,该平台将为DL应用程序提供原生GPU性能。两个核心组件组成了混合无服务器计算架构,一个虚拟GPU (vGPU)层和一个重构容器子系统。shim vGPU层为无服务器并发功能提供高性能GPU共享,具有低延迟和高可扩展性。该层通过使用API远程技术拦截来自无服务器函数的GPU调用,提供细粒度的GPU资源配置和性能隔离。vGPU层通过GPU上下文缓存和位置感知调度来优化无服务器计算中的GPU性能,以减少冷启动和不必要的数据移动。容器子系统通过利用DL模型结构和流水线模型加载来并行处理cpu到gpu的内存复制和模型执行,从而加速了整个DL生命周期。该子系统利用模型分区技术,通过动态地将DL模型分区分配到CPU和GPU,从而加速CPU/GPU混合体系结构。该研究项目设计和实施的科学知识和工具将为下一代云计算和深度学习提供并实现创新。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).Next-generation serverless cloud computing provides developers with simplified access to server management and administration, including event-driven execution, fine-grained resource provisioning, auto-scaling, and pay-as-you-go billing. The machine learning community is taking advantage of these benefits of serverless cloud computing to ease the development and deployment of deep learning (DL) applications. However, existing serverless computing platforms lack efficient support for GPUs, impeding DL practitioners from utilizing serverless computing for large-scale applications. This project will develop an efficient serverless computing platform with a hybrid CPU/GPU architecture to accelerate DL application development and deployment. The goal is to advance cutting-edge methodologies in both deep learning and serverless computing, which will result in a significant leap forward to benefit DL practitioners, DL users, and providers of cloud computing infrastructures, contributing to science advancement for society. The research findings will also enhance undergraduate and graduate education with exciting examples and demonstrations of real-world systems at the intersection of distributed computing, cloud computing, and deep learning.The project will develop a novel serverless computing platform with a hybrid CPU/GPU architecture that will provide DL applications with native GPU performance. Two core components constitute the hybrid serverless computing architecture, a shim virtualized GPU (vGPU) layer and a refactored container subsystem. The shim vGPU layer enables high-performance GPU sharing for concurrent serverless functions with low latency and high scalability. This layer provides fine-grained GPU resource provisioning and performance isolation by intercepting GPU calls from serverless functions using API remoting techniques. The vGPU layer optimizes GPU performance in serverless computing via GPU context caching and locality-aware scheduling to mitigate cold-starts and unnecessary data movement. The container subsystem accelerates the entire DL lifecycle by exploiting DL model structures and pipelined model loading to parallelize CPU-to-GPU memory copy and model execution. The subsystem exploits model partitioning techniques to accelerate the hybrid CPU/GPU architecture by dynamically distributing the DL model partitions to CPU and GPU. The scientific knowledge and tools designed and implemented from this research project will provide and enable innovations for next-generation cloud computing and deep learning.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3485447.3511979
发表时间: 2021-08
期刊: Proceedings of the ACM Web Conference 2022
影响因子: --
作者: [Hanfei Yu;Hao Wang;Jian Li;Xuemei Yuan;Seung-Jong Park]
通讯作者: Hanfei Yu;Hao Wang;Jian Li;Xuemei Yuan;Seung-Jong Park
DOI: 10.1145/3588195.3592996
发表时间: 2023-08
期刊: Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing
影响因子: --
作者: [Hanfei Yu;Christian Fontenot;Hao Wang;Jian Li;Xu Yuan;Seung-Jong Park]
通讯作者: Hanfei Yu;Christian Fontenot;Hao Wang;Jian Li;Xu Yuan;Seung-Jong Park
RII Track-4:NSF: Federated Analytics Systems with Fine-grained Knowledge Comprehension: Achieving Accuracy with Privacy
  • 批准号:
    2327480
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2024
  • 负责人:
    Hao Wang
  • 依托单位:
Collaborative Research: OAC: Core: Harvesting Idle Resources Safely and Timely for Large-scale AI Applications in High-Performance Computing Systems
  • 批准号:
    2403398
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2024
  • 负责人:
    Hao Wang
  • 依托单位:
Collaborative Research: SaTC: CORE: Small: Critical Learning Periods Augmented Robust Federated Learning
  • 批准号:
    2315612
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2023
  • 负责人:
    Hao Wang
  • 依托单位:
RI: Small: Enabling Interpretable AI via Bayesian Deep Learning
  • 批准号:
    2127918
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2021
  • 负责人:
    Hao Wang
  • 依托单位:
国内基金
海外基金
Z8-12:OH和Z8-14:OAc分别维持梨小食心虫和李小食心虫性诱剂特异性的分子基础
  • 批准号:
    --
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    35万元
  • 批准年份:
    2021
  • 负责人:
    陈秀琳
  • 依托单位:
亚硝酰钌配合物[Ru(OAc)(2mqn)2NO]的光异构反应机理研究
  • 批准号:
    21603131
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    19.0万元
  • 批准年份:
    2016
  • 负责人:
    王建茹
  • 依托单位:
机械化学条件下Mn(OAc)3促进的自由基串联反应研究
  • 批准号:
    21242013
  • 项目类别:
    专项基金项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2012
  • 负责人:
    张泽
  • 依托单位: