课题基金 / 基金详情

CAREER: Efficient Large Language Model Inference Through Codesign: Adaptable Software Partitioning and FPGA-based Distributed Hardware

CAREER: Efficient Large Language Model Inference Through Codesign: Adaptable Software Partitioning and FPGA-based Distributed Hardware
职业:通过协同设计进行高效的大型语言模型推理:适应性软件分区和基于 FPGA 的分布式硬件
批准号:
2339084
负责人:
Mohamed Abdelfattah
金额:
$88.31万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-05-01 至 2029-04-30

项目摘要

项目成果

Mohamed Abdelfattah的其他基金

相似基金

相关文献

中文摘要
翻译
人工智能(AI)已经进入了“规模时代”。大量的训练数据被用于在大型计算机上训练巨大的深度神经网络(dnn),大型语言模型(llm)的兴起就是一个缩影。对这项技术的极高需求是显而易见的,最近的一个例子就是ChatGPT:一个LLM聊天机器人,在发布后仅仅两个月就获得了1亿活跃用户,创造了新的世界纪录。然而,部署llm的成本可能相当高,因为它们的内存占用可以扩展到tb级的数据,同时还需要大量的计算资源。因此,大规模分布式计算机已经变得必不可少,特别是为了满足交互式应用程序所要求的性能。为了提高效率,该项目解决了llm特有的新挑战,包括它们的大内存占用、不同的计算需求和分布式计算。这对于使法学硕士更容易获得和可持续地广泛使用至关重要。同时,该奖项旨在通过公立大学针对不同学生群体的大规模人工智能课程、全面的课程整合以及研究生和本科生的学生指导,培养精通算法、硬件和软件的多样化人工智能劳动力。该项目将实现llm和分布式计算平台的协同设计,分为三个主要方向,对应于计算堆栈的三个层次:软件、硬件和算法。最初,该项目将专注于自动分区和映射算法,因为这些构成了llm可以在现有和新的分布式计算平台上部署和优化的基础。这项研究的关键是开发一个可扩展的硬件性能评估器,它可以模拟当前基于gpu的系统以及新的分布式计算方法。特别是,第二个推力研究了分布式系统中使用网络内和近存储fpga来加速LLM推理。最后的推力研究了llm的平台感知压缩,包括混合精度量化和低秩近似。除了提高整个计算栈的LLM效率外,该项目还将开发一个研究框架,以协同优化LLM和分布式硬件平台,从而产生新的优化LLM计算系统和实施方法。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Artificial intelligence (AI) has entered the "age of scale". Huge amounts of training data are being used to train enormous deep neural networks (DNNs) on large-scale computers as epitomized by the rise of large language models (LLMs). The extremely high demand for this technology is clearly evident, as recently exemplified by ChatGPT: an LLM chatbot that garnered 100 million active users merely two months post-release, setting a new world record. However, deploying LLMs can be quite costly, given that their memory footprint can extend to terabytes of data while also demanding high computational resources. Consequently, large-scale distributed computers have become essential, particularly to meet the performance required for interactive applications. To improve efficiency, this project tackles new challenges that are specific to LLMs, including their large memory footprint, varying computational demands, and distributed computing. This is critical to make LLMs more accessible and sustainable for widespread use. Concurrently, this award seeks to develop a diverse AI workforce proficient in algorithms, hardware, and software, achieved through a large-scale AI course for diverse student population at public universities, comprehensive curriculum integration, and student mentorship at both graduate and undergraduate levels.This project will enable the codesign of LLMs and distributed computing platforms, divided into three major thrusts that correspond to three levels of the computing stack: software, hardware, and algorithms. Initially, the project will focus on automated partitioning and mapping algorithms, as these form the foundations by which LLMs can be deployed and optimized on both existing and new distributed computing platforms. Key to this research thrust is the development of an extensible hardware performance estimator that can model current GPU-based systems alongside new distributed computing approaches. In particular, the second thrust investigates the use of in-network and near-storage FPGAs within distributed systems to speed up LLM inference. The final thrust investigates platform-aware compression for LLMs, including mixed-precision quantization and low-rank approximation. In addition to improving LLM efficiency across the computing stack, this project will develop a research framework to synergistically co-optimize LLMs and distributed hardware platforms, resulting in new optimized LLM computing systems and implementation methodologies.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF: Small: Domain-Specific FPGAs to Accelerate Unrolled DNNs with Fine-Grained Unstructured Sparsity and Mixed Precision
  • 批准号:
    2303626
  • 项目类别:
    Standard Grant
  • 资助金额:
    $59.84万
  • 财政年份:
    2023
  • 负责人:
    Mohamed Abdelfattah
  • 依托单位:
海外基金