SHF: Small: Domain-Specific FPGAs to Accelerate Unrolled DNNs with Fine-Grained Unstructured Sparsity and Mixed Precision
SHF: Small: Domain-Specific FPGAs to Accelerate Unrolled DNNs with Fine-Grained Unstructured Sparsity and Mixed Precision
批准号:
2303626
负责人:
Mohamed Abdelfattah
金额:
$59.84万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-05-15 至 2026-04-30
中文摘要
人工智能(AI)已经成为我们日常生活中不可或缺的一部分,它给各个行业带来了革命性的变化,改变了我们与技术互动的方式。人工智能取得显著进步的关键因素之一是深度神经网络(DNN)的效率。这些复杂的系统类似于人脑,擅长处理海量数据,使它们能够学习并做出明智的决定。与传统的计算机程序相比,DNN在图像识别、自然语言处理和决策等任务中表现出了优越的性能。不幸的是,这种改进的性能需要更多的能量和计算资源。这不仅增加了它们的运行成本,还限制了它们在电池供电设备等资源受限环境中的部署,阻碍了人工智能系统的更广泛采用。有趣的是,DNN计算经常涉及许多冗余操作,通常称为“稀疏性”。该项目旨在开发专门的计算机芯片和软件程序,利用丰富的细粒度稀疏性来提高人工智能性能,同时降低能源消耗和计算成本。这一研究奖的成果将被纳入研究生和本科生的教育课程和研究指导计划,以教育下一代计算机工程师,使他们了解硬件/软件共同设计对深度学习的重要性。此外,计划开展一项外联活动,以增加妇女参与人工智能硬件开发。该项目侧重于细粒度非结构化稀疏性和混合精确度的DNN的硬件加速,这两种形式的冗余尚未被现有的计算机芯片有效利用。研究团队专注于在可编程硬件上优化展开的DNN电路,从目前通用的现场可编程门阵列(FGA)硬件结构开始,向DNN优化结构发展。系统基准驱动的方法被用来专门化用于实现展开的DNN电路的FPGA组件。此外,该团队还研究了对FPGA结构的更重大更改,如时分多路复用和内存计算,以增加逻辑容量并支持部署更大的DNN。为了提取最大的效率,共同设计了DNN稀疏算法,包括剪枝、量化和参数共享。这一奖项预计将产生新的位可编程硬件架构、DNN稀疏算法,以及协同设计稀疏DNN和硬件的研究框架。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Artificial intelligence (AI) has become an essential part of our daily lives, revolutionizing various industries and transforming the way we interact with technology. One of the key factors behind AI's remarkable progress is the efficiency of deep neural networks (DNNs). These complex systems, akin to the human brain, excel at processing vast amounts of data, enabling them to learn and make informed decisions. Compared to traditional computer programs, DNNs have shown superior performance in tasks such as image recognition, natural language processing, and decision-making. Unfortunately, this improved performance requires substantially more energy and computational resources. This not only increases their running costs but also limits their deployment in resource-constrained environments like battery-powered devices, hindering the broader adoption of AI systems. Interestingly, DNN computations often involve many redundant operations, termed generally as "sparsity." This project aims to develop specialized computer chips and software programs that exploit abundant fine-grained sparsity to enhance AI performance while reducing energy consumption and computational costs. Outcomes of this research award will be integrated into educational curricula and research mentorship plans at the graduate and undergraduate level, to educate the next generation of computer engineers on the importance of hardware/software codesign for deep learning. In addition, an outreach activity is planned to increase the participation of women in the hardware development for AI.This project focuses on the hardware acceleration of DNNs with fine-grained unstructured sparsity and mixed precision, two forms of redundancy that have yet to be exploited efficiently by existing computer chips. The research team focuses on optimizing unrolled DNN circuits on programmable hardware, starting with the current general-purpose hardware fabric of field-programmable gate arrays (FPGAs) and progressing towards DNN-optimized fabrics. A systematic benchmark-driven approach is used to specialize FPGA components for the implementation of unrolled DNN circuits. Furthermore, the team investigates more significant changes to the FPGA fabric, such as time-multiplexing and in-memory computing, to increase logic capacity and enable the deployment of larger DNNs. To extract maximum efficiency, DNN sparsification algorithms are codesigned, including pruning, quantization, and parameter sharing. This award is expected to result in new bit-programmable hardware architectures, DNN sparsification algorithms, and a research framework to synergistically codesign sparse DNNs and hardware.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
BRAMAC: Compute-in-BRAM Architectures for Multiply-Accumulate on FPGAs
BRAMAC:用于 FPGA 上乘法累加的 BRAM 计算架构
DOI:
10.1109/fccm57271.2023.00015
发表时间:
2023
期刊:
Proceedings Annual IEEE Symposium on Field Programmable Custom Computing Machines
影响因子:
--
作者:
[Chen, Yuzong, Abdelfattah, Mohamed S.]
通讯作者:
Abdelfattah, Mohamed S.
DOI:
10.1109/icfpt59805.2023.00013
发表时间:
2023-11
期刊:
2023 International Conference on Field Programmable Technology (ICFPT)
影响因子:
--
作者:
[Yuzong Chen;Jordan Dotzel;M. Abdelfattah]
通讯作者:
Yuzong Chen;Jordan Dotzel;M. Abdelfattah
CAREER: Efficient Large Language Model Inference Through Codesign: Adaptable Software Partitioning and FPGA-based Distributed Hardware
-
批准号:2339084
-
项目类别:Continuing Grant
-
资助金额:$88.31万
-
财政年份:2024
-
负责人:Mohamed Abdelfattah
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: