Exploiting and Enhancing Programmable Logic for Deep Learning and Datacenter Acceleration
Exploiting and Enhancing Programmable Logic for Deep Learning and Datacenter Acceleration
批准号:
RGPIN-2022-04445
负责人:
Betz, Vaughn
金额:
$6.63万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
计算改变了社会,实现了无处不在的通信、点播娱乐、语音识别等。虽然深度学习(DL)和其他应用的计算需求正在增长,但晶体管扩展带来的传统效率增长正在放缓。为了填补这一空白,我们需要更高效但可重新编程的计算设备来支持新的应用。现场可编程门阵列(FGA)可以在硬件级别进行重新编程,使许多嵌入式和数据中心应用的能效比处理器高出10倍或更多。我们寻求在三个相关方面取得进展:在现场可编程门阵列上实现高效的DL推理,设计更好的可重构器件,以及增强计算机辅助设计(CAD)工具以支持这些新器件。我们的第一个研究重点是在创建生产性开发流的同时,寻求对现场可编程门阵列进行更高效的DL推理。在我们的异质流水线(HPIPE)项目中,我们通过使用新的域特定编译器为卷积神经网络(CNN)中的每一层实现定制硬件,从而利用了FPGA的可编程性。我们的神经处理单元(NPU)项目改为创建由软件产生的指令流控制的DL功能单元。这两个项目都具有行业领先的性能,我们将通过多种方式增强它们,包括并行扩展到多个芯片,利用AI优化的现场可编程门阵列中的新张量块,以及将HPIPE的专用单元与NPU的软件可编程性相结合。我们的第二个目标是寻求新的可重新配置加速器设备(RAD)架构,以实现更高的性能和更容易的开发,特别是针对DL和数据中心基础设施。我们不仅将研究传统的2D芯片,还将研究由最新技术实现的多芯片堆栈。我们设想的RAD将在包含粗粒度可编程加速器(如强化的矩阵-矢量乘法单元)、大容量存储块和用于链接所有组件的嵌入式片上网络(NOC)的基础架构芯片上结合一个FPGA结构芯片。FPGA结构和粗粒度加速器的结合可以提高性能,而NoC将设计组件解耦以简化设计。我们的第三个推力是开发计算机辅助设计(CAD)工具来研究这些RAD架构,并允许在它们上实施DL应用程序。首先,我们将开发一个新工具(RADSim),通过确定每个架构上各种应用程序的执行时间来评估交换矩阵、加速器和NoC组合。接下来,我们将增强广泛使用的通用位置和路由(VPR)工具,以共同优化交换矩阵资源、加速器块和NoC路由器的放置,并通过RADsim提供延迟和拥塞估计。开源的vpr工具已经支持了各种各样的创新和产品,这些增强将使它更有能力。
英文摘要
Computing has transformed society, enabling ubiquitous communications, on-demand entertainment, speech recognition, and more. While computation demand is growing for deep learning (DL) and other applications, the traditional efficiency gains afforded by transistor scaling are slowing down. To fill this gap, we need computational devices that are more efficient yet reprogrammable to allow new applications. Field-Programmable Gate Arrays (FPGAs) can be reprogrammed at the hardware level, enabling energy-efficiency gains of 10x or more vs. processors for many embedded and datacenter applications. We seek to advance on three related fronts: implementing efficient DL inference on FPGAs, architecting better reconfigurable devices and enhancing computer-aided design (CAD) tools to enable these new devices. Our first research thrust seeks more efficient DL inference on FPGAs while simultaneously creating productive development flows. In our heterogeneous pipeline (HPIPE) project, we leverage FPGA programmability by implementing customized hardware for every layer in a convolutional neural network (CNN) using a new domain specific compiler. Our neural processing unit (NPU) project instead creates DL functional units controlled by an instruction stream produced from software. Both projects have industry-leading performance, and we will enhance them in multiple ways, including scaling to multiple chips in parallel, exploiting the new tensor blocks in AI-optimized FPGAs, and combining the specialized units of HPIPE with the software programmability of the NPU. Our second thrust seeks new reconfigurable accelerator device (RAD) architectures to allow higher performance and easier development, particularly for DL and for datacenter infrastructure. We will investigate not only conventional 2D chips, but also the multi-die stacks enabled by recent technologies. We envision RADs that combine an FPGA fabric die on an infrastructure die containing coarse-grain programmable accelerators (such as hardened matrix-vector multiply units), large memory blocks, and an embedded network-on-chip (NoC) to link all the components. The combination of FPGA fabric and coarse-grained accelerators can increase performance, while the NoC decouples design components to simplify design. Our third thrust develops the computer-aided design (CAD) tools to investigate these RAD architectures and allow implementation of DL applications on them. First, we will develop a new tool (RADSim) to evaluate fabric, accelerator and NoC combinations by determining execution time for various applications on each architecture. Next, we will enhance the widely-used Versatile Place and Route (VPR) tool to co-optimize the placement of fabric resources, accelerator blocks and NoC routers, with latency and congestion estimates informed by RADsim. The open-source VPR tool is already enabling a wide variety of innovation and products, and these enhancements will make it still more capable.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Toward More Energy-Efficient Datacenters with Enhanced Programmable Silicon
-
批准号:RGPIN-2016-05537
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.93万
-
财政年份:2021
-
负责人:Betz, Vaughn
-
依托单位:
Toward More Energy-Efficient Datacenters with Enhanced Programmable Silicon
-
批准号:RGPIN-2016-05537
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.93万
-
财政年份:2020
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Intel Industrial Research Chair in Programmable Silicon
-
批准号:428842-2016
-
项目类别:Industrial Research Chairs
-
资助金额:$17.48万
-
财政年份:2020
-
负责人:Betz, Vaughn
-
依托单位:
Toward More Energy-Efficient Datacenters with Enhanced Programmable Silicon
-
批准号:RGPIN-2016-05537
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.93万
-
财政年份:2019
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Intel Industrial Research Chair in Programmable Silicon
-
批准号:428842-2016
-
项目类别:Industrial Research Chairs
-
资助金额:$4.69万
-
财政年份:2019
-
负责人:Betz, Vaughn
-
依托单位:
Toward More Energy-Efficient Datacenters with Enhanced Programmable Silicon
-
批准号:RGPIN-2016-05537
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.93万
-
财政年份:2018
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Intel Industrial Research Chair in Programmable Silicon
-
批准号:428842-2016
-
项目类别:Industrial Research Chairs
-
资助金额:$8.74万
-
财政年份:2018
-
负责人:Betz, Vaughn
-
依托单位:
Toward More Energy-Efficient Datacenters with Enhanced Programmable Silicon
-
批准号:RGPIN-2016-05537
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.93万
-
财政年份:2017
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Intel Industrial Research Chair in Programmable Silicon
-
批准号:428842-2016
-
项目类别:Industrial Research Chairs
-
资助金额:$8.74万
-
财政年份:2017
-
负责人:Betz, Vaughn
-
依托单位:
Fast and accurate biophotonic simulations for photodynamic cancer therapy treatment planning
-
批准号:490784-2015
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$5.2万
-
财政年份:2017
-
负责人:Betz, Vaughn
-
依托单位:
Toward More Energy-Efficient Datacenters with Enhanced Programmable Silicon
-
批准号:RGPIN-2016-05537
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.93万
-
财政年份:2016
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Altera Industrial Research Chair in Programmable Silicon
-
批准号:418003-2010
-
项目类别:Industrial Research Chairs
-
资助金额:$7.47万
-
财政年份:2015
-
负责人:Betz, Vaughn
-
依托单位:
Next-generation programmable logic devices: improving silicon efficiency and designer productivity
-
批准号:403299-2011
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2015
-
负责人:Betz, Vaughn
-
依托单位:
Next-generation programmable logic devices: improving silicon efficiency and designer productivity
-
批准号:403299-2011
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2014
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Altera Industrial Research Chair in Programmable Silicon
-
批准号:418003-2010
-
项目类别:Industrial Research Chairs
-
资助金额:$9.65万
-
财政年份:2014
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Altera Industrial Research Chair in Programmable Silicon
-
批准号:418003-2010
-
项目类别:Industrial Research Chairs
-
资助金额:$8.56万
-
财政年份:2013
-
负责人:Betz, Vaughn
-
依托单位:
Next-generation programmable logic devices: improving silicon efficiency and designer productivity
-
批准号:403299-2011
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2013
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Altera Industrial Research Chair in Programmable Silicon
-
批准号:418003-2010
-
项目类别:Industrial Research Chairs
-
资助金额:$9.65万
-
财政年份:2012
-
负责人:Betz, Vaughn
-
依托单位:
Next-generation programmable logic devices: improving silicon efficiency and designer productivity
-
批准号:403299-2011
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2012
-
负责人:Betz, Vaughn
-
依托单位:
NSERC/Altera Industrial Research Chair in Programmable Silicon
-
批准号:418003-2010
-
项目类别:Industrial Research Chairs
-
资助金额:$8.38万
-
财政年份:2011
-
负责人:Betz, Vaughn
-
依托单位:
海外基金