课题基金 / 基金详情

CRII: CNS: System for Deploying Ultra Low-Latency Machine Learning Applications on Programmable Networks

CRII: CNS: System for Deploying Ultra Low-Latency Machine Learning Applications on Programmable Networks
CRII:CNS:在可编程网络上部署超低延迟机器学习应用程序的系统
批准号:
2245352
负责人:
Sean Choi
金额:
$17.42万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2025-09-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
许多现代应用,如自动驾驶、安全威胁检测和图像识别,都依赖于机器学习(ML)模型,这些模型是使用大量数据构建的统计模型,可以自动实现一系列复杂问题的解决方案。ML模型通常非常复杂,因此它们通常需要非常强大的计算硬件和大量的时间(通常需要几分钟)来获得解决方案。然而,包括上面提到的那些应用程序在内的现代应用程序需要在几毫秒内做出决策,以便能够对环境中的变化做出反应。因此,机器学习的一个主要挑战是开发方法和计算机系统,使机器学习模型能够非常快速地为复杂问题提供解决方案,同时最大限度地减少模型所需的硬件数量。解决这样的挑战不仅会增加在现代应用中使用复杂ML模型的可行性,而且还会为提高我们国家安全的国防系统提供潜在的改进。此外,这项研究直接用于开发新的计算机系统课程,并为许多本科生提供了参与研究的机会,其中许多人是第一次参与研究。该项目专注于解决在显著减少时间(定义为延迟)方面的挑战,为使用ML模型的应用程序提供解决方案。该项目将探索的一种方法是将传统上运行在中央处理单元(CPU)或图形处理单元(GPU)上的模型运行在称为网络处理单元(NPU)的特定领域架构(DSA)上。npu是存在于网络设备(如网卡)中的计算硬件,是进出计算机的数据的网关。使用npu的主要动机是减轻将数据传递到CPU或gpu的开销,从而减少处理数据的延迟、CPU周期和内存,同时提供性能保证,减少对上下文切换的需求。这种方法面临的一系列挑战是:(1)确定可以通过可衡量的改进卸载到npu上的ML应用程序的类型;(2)以高效和可扩展的方式在npu上编程和部署应用程序;(3)利用现有流量保证可预测的性能。该项目将专注于开发将机器学习应用程序部署到npu并量化其效益的方法和系统,重点关注现有的、更简单的机器学习模型,如决策树和逻辑回归,以显示其方法的可行性,并获得性能改进的初步指标。整个工作将作为开源、可重用的库和应用程序发布,供其他研究人员使用。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many modern applications, such as self-driving, security-threat detection, and image recognition rely on machine learning (ML) models, which are statistical models built using large amounts of data to automatically achieve solutions for a set of complex problems. ML models are often very complex, thus they often require very powerful computing hardware and large amounts of time, often taking minutes, to arrive at a solution. However, modern applications including those mentioned above require decisions to be made in milliseconds to be able to react to the changes in the environment. Therefore, a major challenge in machine learning is to develop methods and computer systems that allow ML models to be able to provide solutions to complex problems very quickly, while minimizing the amount of hardware that the models need. Solving such a challenge will not only increase the feasibility of using complex ML models for modern applications, but also, for example, provide potential improvements for defense systems that improve our national security. Furthermore, this research directly feeds into the development of new computer systems courses and provides opportunities for a number of undergraduates—many for the first time—to participate in research.This project focuses on solving the challenges in providing significant reduction in time—defined as latency—to provide solutions for applications that use ML models. An approach that will be explored by this project is for models that traditionally are run on a central processing unit (CPU) or a graphics processing unit (GPU) to run on a domain-specific architecture (DSA) called a network processing unit (NPU). NPUs are computational hardware that exist in networking devices, such as network interface cards (NICs), and act as the gateway to data that enters and leaves a computer. The main motivation for using NPUs is to mitigate the overhead of passing data to CPUs or GPUs, thereby reducing the latency, CPU cycles and memory spent on processing the data, while providing performance guarantees that come with the reduced need for context switching. The set of challenges for this approach are: (1) determining the types of ML applications that are feasible to be offloaded onto NPUs with measurable improvements; (2) programming and deploying applications on NPUs in an efficient and scalable manner; and (3) guaranteeing predictable performance with existing traffic. This project will focus on developing methods and a system for deploying ML applications on to NPUs and quantifying their benefits, focusing on existing, simpler ML models, such as decision trees and logistic regression, to show feasibility of its approach and to obtain preliminary metrics on performance improvements. The entire work will be released as open-source, reusable libraries, and applications for use by other researchers.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
IL-17A通过STAT5影响CNS2区域甲基化抑制调节性T细胞功能在银屑病发病中的作用和机制研究
miR-20a通过调控CD4+T细胞焦亡促进CNS炎性脱髓鞘疾病的发生及机制研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    王亦舒
  • 依托单位:
血浆CNS来源外泌体中寡聚磷酸化α-synuclein对PD病程的提示研究
  • 批准号:
    82101506
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    徐妍
  • 依托单位:
基于脑微血管内皮细胞模型的毒力岛4在单增李斯特菌CNS炎症中的作用及机制研究
  • 批准号:
    32160834
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    35万元
  • 批准年份:
    2021
  • 负责人:
    马勋
  • 依托单位: