CRII: CNS: System for Deploying Ultra Low-Latency Machine Learning Applications on Programmable Networks
CRII: CNS: System for Deploying Ultra Low-Latency Machine Learning Applications on Programmable Networks
批准号:
2245352
负责人:
Sean Choi
金额:
$17.42万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2025-09-30
中文摘要
许多现代应用,如自动驾驶、安全威胁检测和图像识别,都依赖于机器学习(ML)模型,这些模型是使用大量数据构建的统计模型,可以自动实现一系列复杂问题的解决方案。ML模型通常非常复杂,因此它们通常需要非常强大的计算硬件和大量的时间(通常需要几分钟)来获得解决方案。然而,包括上面提到的那些应用程序在内的现代应用程序需要在几毫秒内做出决策,以便能够对环境中的变化做出反应。因此,机器学习的一个主要挑战是开发方法和计算机系统,使机器学习模型能够非常快速地为复杂问题提供解决方案,同时最大限度地减少模型所需的硬件数量。解决这样的挑战不仅会增加在现代应用中使用复杂ML模型的可行性,而且还会为提高我们国家安全的国防系统提供潜在的改进。此外,这项研究直接用于开发新的计算机系统课程,并为许多本科生提供了参与研究的机会,其中许多人是第一次参与研究。该项目专注于解决在显著减少时间(定义为延迟)方面的挑战,为使用ML模型的应用程序提供解决方案。该项目将探索的一种方法是将传统上运行在中央处理单元(CPU)或图形处理单元(GPU)上的模型运行在称为网络处理单元(NPU)的特定领域架构(DSA)上。npu是存在于网络设备(如网卡)中的计算硬件,是进出计算机的数据的网关。使用npu的主要动机是减轻将数据传递到CPU或gpu的开销,从而减少处理数据的延迟、CPU周期和内存,同时提供性能保证,减少对上下文切换的需求。这种方法面临的一系列挑战是:(1)确定可以通过可衡量的改进卸载到npu上的ML应用程序的类型;(2)以高效和可扩展的方式在npu上编程和部署应用程序;(3)利用现有流量保证可预测的性能。该项目将专注于开发将机器学习应用程序部署到npu并量化其效益的方法和系统,重点关注现有的、更简单的机器学习模型,如决策树和逻辑回归,以显示其方法的可行性,并获得性能改进的初步指标。整个工作将作为开源、可重用的库和应用程序发布,供其他研究人员使用。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many modern applications, such as self-driving, security-threat detection, and image recognition rely on machine learning (ML) models, which are statistical models built using large amounts of data to automatically achieve solutions for a set of complex problems. ML models are often very complex, thus they often require very powerful computing hardware and large amounts of time, often taking minutes, to arrive at a solution. However, modern applications including those mentioned above require decisions to be made in milliseconds to be able to react to the changes in the environment. Therefore, a major challenge in machine learning is to develop methods and computer systems that allow ML models to be able to provide solutions to complex problems very quickly, while minimizing the amount of hardware that the models need. Solving such a challenge will not only increase the feasibility of using complex ML models for modern applications, but also, for example, provide potential improvements for defense systems that improve our national security. Furthermore, this research directly feeds into the development of new computer systems courses and provides opportunities for a number of undergraduates—many for the first time—to participate in research.This project focuses on solving the challenges in providing significant reduction in time—defined as latency—to provide solutions for applications that use ML models. An approach that will be explored by this project is for models that traditionally are run on a central processing unit (CPU) or a graphics processing unit (GPU) to run on a domain-specific architecture (DSA) called a network processing unit (NPU). NPUs are computational hardware that exist in networking devices, such as network interface cards (NICs), and act as the gateway to data that enters and leaves a computer. The main motivation for using NPUs is to mitigate the overhead of passing data to CPUs or GPUs, thereby reducing the latency, CPU cycles and memory spent on processing the data, while providing performance guarantees that come with the reduced need for context switching. The set of challenges for this approach are: (1) determining the types of ML applications that are feasible to be offloaded onto NPUs with measurable improvements; (2) programming and deploying applications on NPUs in an efficient and scalable manner; and (3) guaranteeing predictable performance with existing traffic. This project will focus on developing methods and a system for deploying ML applications on to NPUs and quantifying their benefits, focusing on existing, simpler ML models, such as decision trees and logistic regression, to show feasibility of its approach and to obtain preliminary metrics on performance improvements. The entire work will be released as open-source, reusable libraries, and applications for use by other researchers.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
IL-17A通过STAT5影响CNS2区域甲基化抑制调节性T细胞功能在银屑病发病中的作用和机制研究
-
批准号:82304006
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:刘阳禾
-
依托单位:
miR-20a通过调控CD4+T细胞焦亡促进CNS炎性脱髓鞘疾病的发生及机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:王亦舒
-
依托单位:
血浆CNS来源外泌体中寡聚磷酸化α-synuclein对PD病程的提示研究
-
批准号:82101506
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:徐妍
-
依托单位:
基于脑微血管内皮细胞模型的毒力岛4在单增李斯特菌CNS炎症中的作用及机制研究
-
批准号:32160834
-
项目类别:地区科学基金项目
-
资助金额:35万元
-
批准年份:2021
-
负责人:马勋
-
依托单位:
胱硫醚-β-合成酶介导小胶质细胞极化致糖皮质激素CNS毒性作用及机制研究
-
批准号:82104317
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:赵瑛
-
依托单位:
生物工程化微泡干扰MAPK通路重编程CNS微环境起始脑胶质瘤免疫检查点抑制剂的应答研究
-
批准号:82102900
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:卢利森
-
依托单位:
百合中“桉树脑盒”挥发物生物合成关键酶基因CNS的功能解析
-
批准号:32002082
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:孔滢
-
依托单位:
新型化合物组合抑制STAT6维持Foxp3-CNS2去甲基化产生稳定的iTreg细胞诱导小鼠肾移植免疫耐受的机制研究
-
批准号:82070773
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2020
-
负责人:陈恕求
-
依托单位:
大气细颗粒物通过NF-κB/LBP-9信号通路诱导小胶质细胞激活加剧CNS脱髓鞘损伤的作用机制研究
-
批准号:82071396
-
项目类别:面上项目
-
资助金额:55.0万元
-
批准年份:2020
-
负责人:张媛
-
依托单位:
环状RNA介导CNS1S1基因影响热应激奶牛乳腺αs1-casein合成及机制研究
-
批准号:32072714
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2020
-
负责人:孙加节
-
依托单位: