EDGE - Adaptive Deep Learning Hardware for Embedded Platforms
EDGE - Adaptive Deep Learning Hardware for Embedded Platforms
批准号:
EP/V034111/1
负责人:
Xiaojun Zhai
金额:
$29.58万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
深度学习(DL)是现代人工智能(AI)的关键技术,它为许多基于机器学习的应用提供了最先进的准确性。如今,尽管DL系统的大部分计算负载仍然花费在数据中心运行神经网络上,但智能手机的普及以及用于增强现实(AR)、虚拟现实(VR)和自主机器人系统的独立可穿戴设备的即将推出,对具有高能量和计算效率的DL推理硬件提出了很高的要求,沿着DL技术的快速发展。最近,我们见证了DL架构类型的明显演变,提出了更复杂的网络架构来改善边缘AI推理。这包括动态网络架构,它以数据依赖的方式随每个新输入而变化,其中输入和内部状态不固定。DL中的这种新架构概念很可能会影响未来交付此类功能所需的硬件架构类型。该项目正是针对这一挑战,并提出设计一个灵活的硬件架构,使嵌入式设备上的各种DL算法的自适应支持。首先,为了生产可以针对专用应用规范和操作环境自适应地配置的更低成本、更低功率和更高处理效率的DL推理硬件,这将需要在当前DL技术的软件和硬件的优化方面进行彻底的创新。开发和实践演示器,以实现对嵌入式边缘设备上的各种DL技术的一般支持,资源和延迟预算有限。首先,这需要在计算体系结构、存储器层次结构和资源利用率以及系统延迟和吞吐量方面对当前DL体系结构进行彻底的创新:对于现代DL系统而言,推理过程是动态的,例如DL推理可能依赖于输入和资源。因此,该提案寻求以下三个重点:首先,在PI现有工作的基础上,为资源受限的嵌入式设备优化机器学习模型,以实现网络模型可以根据需要通过硬件感知近似技术动态优化的目标。第二,利用新开发的可编程存储器层次结构和自适应处理硬件中的自适应计算加速技术,寻求一个新的雄心勃勃的方向,开发一套上下文感知硬件架构,与能够充分利用真正硬件能力的近似算法密切合作。与DL硬件推理引擎的传统优化技术不同,拟议的工作将探索自适应计算加速技术的软件和硬件可编程性,以最大限度地提高目标应用场景的优化结果。第三,该项目将与我们的行业和项目合作伙伴密切合作,制作一个实用的演示器,以展示拟议的DL框架与传统方法相比的有效性,特别是评估该框架在现实世界的关键任务应用中的有效性。
英文摘要
Deep learning (DL) is the key technique in modern artificial intelligence (AI), which has provided state-of-the-art accuracy on many machine-learning based applications. Today, although most of the computational loads of DL systems are still spent running neural networks in data centres, the ubiquity of smartphones, and the upcoming availability of self-contained wearable devices for augmented reality (AR), virtual reality (VR) and autonomous robot systems are placing heavy demands on DL-inference hardware with high energy and computing efficiencies along with rapid development of DL techniques. Recently, we have witnessed a distinct evolution in the types of DL architecture, with more sophisticated network architectures proposed to improve edge AI inference. This includes dynamic network architectures that change with each new input in a data-dependent way, where inputs and internal states are not fixed. Such new architectural concepts in DL are likely to affect the type of hardware architectures that will be required to deliver such capabilities in the future. This project precisely addresses this challenge and proposes to design a flexible hardware architecture that enables adaptive support for a variety of DL algorithms on embedded devices. Primarily, to produce lower cost, lower power and higher processing efficiency DL-inference hardware that can be configured adaptably for dedicated application specifications and operating environments, this will require radical innovation in the optimisation of both the software and the hardware of current DL techniques.This work aims to perform fundamental research, development and practical demonstrator to enable general support for a variety of DL techniques on embedded edge devices with limited resource and latency budgets. Primarily, this requires radical innovation on the current DL architectures in terms of computing architecture, memory hierarchy and resource utilisation, as well as system latency and throughput: it is particularly important for the modern DL systems that the inference processes are dynamic, such as, the DL inference maybe input-dependent and resource-dependent. The proposal therefore seeks the following three thrusts: First, to build upon the existing work of the PI in optimising machine-learning models for resource-constrained embedded devices, towards achieving the goal that the network model could be dynamically optimised as needed through hardware-aware approximation techniques. Second, with newly-developed adaptive compute acceleration technology in programmable memory hierarchy and adaptive processing hardware, to seek a new ambitious direction to develop a set of context-aware hardware architectures to work closely with the approximation algorithms that can fully utilise the true hardware capabilities. Unlike traditional optimisation techniques for DL hardware inference engines, the proposed work will explore both software and hardware programmability of adaptive compute acceleration technology, in order to maximise the optimisation results for the target application scenarios. Third, this project will work closely with our industry and project partners to produce a practical demonstrator to showcase the effectiveness of the proposed DL framework versus traditional approaches, particularly, evaluating the effectiveness of the framework in real-world mission-critical applications.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Application Level Resource Scheduling for Deep Learning Acceleration on MPSoC
MPSoC 上深度学习加速的应用级资源调度
DOI:
10.1007/s11265-023-01881-9
发表时间:
2023
期刊:
Journal of Signal Processing Systems
影响因子:
--
作者:
[Gao C]
通讯作者:
Gao C
DOI:
10.1109/newcas57931.2023.10198162
发表时间:
2023-06
期刊:
2023 21st IEEE Interregional NEWCAS Conference (NEWCAS)
影响因子:
--
作者:
[Cong Gao;Xuqi Zhu;S. Saha;K. Mcdonald-Maier;X. Zhai]
通讯作者:
Cong Gao;Xuqi Zhu;S. Saha;K. Mcdonald-Maier;X. Zhai
Anomaly Behaviour tracing of CHERI-RISC V using Hardware-Software Co-design
使用软硬件协同设计的 CHERI-RISC V 异常行为追踪
DOI:
10.1109/newcas57931.2023.10198103
发表时间:
2023
期刊:
影响因子:
--
作者:
[Borowski M]
通讯作者:
Borowski M
DOI:
10.1109/icac55051.2022.9911081
发表时间:
2022-09
期刊:
2022 27th International Conference on Automation and Computing (ICAC)
影响因子:
--
作者:
[Cong Gao;S. Saha;Yufan Lu;Rappy Saha;K. Mcdonald-Maier;X. Zhai]
通讯作者:
Cong Gao;S. Saha;Yufan Lu;Rappy Saha;K. Mcdonald-Maier;X. Zhai
Pattern Recognition and Artificial Intelligence - Third International Conference, ICPRAI 2022, Paris, France, June 1-3, 2022, Proceedings, Part II
模式识别和人工智能 - 第三届国际会议,ICPRAI 2022,法国巴黎,2022 年 6 月 1-3 日,会议记录,第二部分
DOI:
10.1007/978-3-031-09282-4_10
发表时间:
2022
期刊:
影响因子:
--
作者:
[Boukhennoufa I]
通讯作者:
Boukhennoufa I
共 7 条
海外基金