课题基金 / 基金详情

CAREER: Scalable and Adaptable Sparsity-driven Methods for more Efficient AI Systems

CAREER: Scalable and Adaptable Sparsity-driven Methods for more Efficient AI Systems
职业:可扩展且适应性强的稀疏驱动方法,可实现更高效的人工智能系统
批准号:
2238291
负责人:
Gheorghi Guzun
金额:
$55.03万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-03-01 至 2028-02-29

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Artificial Intelligence (AI) and, in particular, Deep Neural Networks (DNN) have achieved better than human accuracy on many cognitive tasks involving images, natural language processing, and protein structure, among others. Unfortunately, due to high data processing demands, AI systems are typically run on power-hungry specialized computing hardware. Quantization, or approximation to smaller numerical values, has been used to reduce computing requirements. However, the fixed low bit-width DNNs may suffer losses in accuracy due to quantization errors. Many existing software solutions for quantization are also fixed or limited in bit-width choices. To address this trade-off and leverage data sparsity, the research team will investigate state-of-the-art methods and develop novel data quantization, encoding, and compression algorithms to integrate with existing AI systems. The methods developed have the potential to not only improve performance but also to reduce power requirements and boost the energy efficiency of AI systems. They will enable AI applications such as DNN inference on small devices, thus reducing the load on cloud infrastructure, improving user experience, providing data privacy, and avoiding security risks. The work proposed in this project has the potential to push the boundaries in many AI applications that run on energy storage-constrained devices, such as smart sensing, wearable devices, and autonomous driving. The research and educational tools will facilitate and increase student and research community participation in advancing AI research. The work will be conducted at a minority-serving institution, and the funding will support students from underrepresented groups.The research goal of this project is to investigate quantization and compression methods that can leverage sparsity and improve efficiency in AI systems. The principal investigator (PI) plans to study adaptable quantization and compression methods to leverage sparsity in AI systems while minimizing the overhead in non-sparse situations and minimizing accuracy loss. The trade-off between accuracy and performance with the proposed methods will be studied and defined for automated tunable prioritization of either accuracy, performance, or energy efficiency. The PI plans to develop a prototype with parallel execution of the proposed methods to make the proposed methods truly effective for data centers and advanced hardware architectures. The proposed methods will be packaged into an AI vector primitives library that will be integrated with several popular Deep Learning frameworks as proof of concept, primarily targeting GPU and CPU systems. An integration API will be developed for frameworks like Pytorch or TensorFlow to allow easy integration with other vector primitives. Software libraries will be integrated with a web-based learning platform with automated feedback and a motivating environment to encourage student participation in solving AI challenges.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis