课题基金 / 基金详情

SHF: Small: Improving Efficiency of Vision Transformers via Software-Hardware Co-Design and Acceleration

SHF: Small: Improving Efficiency of Vision Transformers via Software-Hardware Co-Design and Acceleration
SHF:小型:通过软硬件协同设计和加速提高视觉变压器的效率
批准号:
2233893
负责人:
Avesta Sasan
金额:
$45.01万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2026-09-30

项目摘要

项目成果

Avesta Sasan的其他基金

相似基金

相关文献

中文摘要
翻译
Transformer模型是机器学习领域相对较新的突破,它彻底改变了自然语言处理,并促进了计算机视觉模型的泛化。然而,Transformer模型的广泛采用需要使它们显著地更节能。Transformer模型过于复杂,现有的硬件对于它们的高效执行来说不是最佳的。在这个项目中,研究人员正在探索两个相互关联的研究问题,以解决Transformer的效率:1)创建新的Transformer模型,可以动态修剪,以提高效率,而不牺牲准确性,2)设计专门的硬件,使变压器的执行更有效。该项目的影响在几个方面是显著的:首先,它通过推进社区对注意力的理解促进了研究社区的科学进步,注意力是Transformer模型中的主要机制,以及它如何用于复杂学习模型的上下文感知修剪。其次,它扩展了硬件社区的知识,为复杂和可调硬件系统的动态精度调整和调度设计了高级解决方案。此外,该项目还支持加州大学戴维斯分校的多样性和高等教育,同时通过将研究融入加州大学戴维斯分校的教学来改善教育。最终,该项目的成功将使上级Transformer模型在各种应用中更容易访问,从而造福社会。研究人员在他们的新Transformer模型中探索了一种增量采样方法,以跨编码器层处理输入图像,逐步获得上下文感知。他们的目标是利用增量上下文感知来删除无人值守的令牌,并屏蔽新样本中不重要的输入补丁。此外,研究人员还探索了基于学习和上下文感知的注意力头丢弃、编码器层跳过和粗粒度模型修剪的提前终止。为了提高Transformer模型的推理效率,研究人员探索构建一个随机预处理单元,该单元近似于矩阵-矩阵乘法,支持基于注意力的模型修剪分类器,用于补丁,令牌,注意力头和编码器消除。为了构建硬件加速器的乘法和累加(MAC)单元,研究人员探索了一种新的临时进位延迟解决方案,消除了MAC中的进位传播。该解决方案简化了MAC逻辑,提高了流处理速度和效率。此外,研究人员的目标是利用新MAC的浅逻辑深度来设计高效的可扩散MAC,从而在处理阵列中实现动态精度权衡。研究人员还研究开发一种调度器,用于在稀疏注意力图上操作时平衡工作负载和最小化内存访问,支持无序令牌处理,处理元素聚类,有序完成,和精确度-该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查进行评估,被认为值得支持的搜索.
英文摘要
Transformer models are a relatively recent breakthrough in machine learning that have revolutionized natural language processing and boosted the generalization of computer vision models. However, the wide adoption of transformer models requires making them significantly more energy efficient. The Transformer models are too complex, and existing hardware is not optimal for their efficient execution. In this project, researchers are exploring two interrelated research problems to tackle transformer efficiency: 1) creating new Transformer models that can be dynamically pruned to improve efficiency without sacrificing accuracy, and 2) designing specialized hardware to make the execution of Transformers more efficient. The impact of this project is significant in several ways: Firstly, it promotes scientific progress in the research community by advancing communities understanding of attention as the primary mechanism in transformer models and how it can be used for context-aware pruning of complex learning models. Secondly, it extends knowledge of the hardware community in designing advanced solutions for dynamic precision tuning and scheduling of complex and tunable hardware systems. Additionally, the project supports diversity and higher education at University of California (UC) – Davis while improving education by integrating research into teaching at UC Davis classes. Ultimately, the project's success will benefit society by making superior transformer models more accessible across various applications.Researchers explore an incremental sampling approach in their new transformer model to process input images across encoder layers gaining contextual awareness progressively. They aim to leverage incremental contextual awareness to remove unattended tokens and mask unimportant input patches in new samples. Additionally, researchers explore learning-based and context-aware attention-head dropping, encoder-layer skipping, and early termination for coarse grain model pruning. To improve the transformer model's inference efficiency, the researchers explore architecting a stochastic pre-processing unit that approximates matrix-matrix multiplication supporting attention-based model pruning classifiers for patch, token, attention head, and encoder elimination. To build the hardware accelerator's multiplication and accumulation (MAC) units, researchers explore a novel solution for temporal carry-bit deferment, eliminating carry-bit propagation in MAC. This solution simplifies MAC logic, enhancing stream processing speed and efficiency. Furthermore, Researchers aim to leverage the shallow logic depth of the new MAC to design highly efficient diffusible MACs, enabling dynamic precision trade-offs in the processing array. Researchers also investigate developing a scheduler for balancing workload and minimizing memory accesses when operating on sparse attention graphs with support for out-of-order token processing, processing element clustering, in-order completion, and precision-aware scheduling to optimize performance.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SaTC: STARSS: Small: IoT Circuit Locking, Obfuscation & Authentication Kernel (CLOAK), A Compilable Architecture for Secure IoT Device Production, Testing, Activation & Ope
  • 批准号:
    2200446
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2021
  • 负责人:
    Avesta Sasan
  • 依托单位:
CSR: Small: Evolution of Computer Vision for Low Power Devices, Breaking its Power Wall and Computational Complexity
  • 批准号:
    2146726
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2021
  • 负责人:
    Avesta Sasan
  • 依托单位:
SaTC: STARSS: Small: IoT Circuit Locking, Obfuscation & Authentication Kernel (CLOAK), A Compilable Architecture for Secure IoT Device Production, Testing, Activation & Ope
  • 批准号:
    1718434
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2017
  • 负责人:
    Avesta Sasan
  • 依托单位:
CSR: Small: Evolution of Computer Vision for Low Power Devices, Breaking its Power Wall and Computational Complexity
  • 批准号:
    1718538
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2017
  • 负责人:
    Avesta Sasan
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: