AdaFlow: A Framework for Adaptive Dataflow CNN Acceleration on FPGAs

AdaFlow: A Framework for Adaptive Dataflow CNN Acceleration on FPGAs
复制标题

AdaFlow:FPGA 上的自适应数据流 CNN 加速框架

DOI:
--
复制
发表时间:
2022
期刊:
Design, Automation and Test in Europe
影响因子:
--
通讯作者:
A. C. S. Beck
A. C. S. Beck
中科院分区:
--
文献类型:
--
作者:
Guilherme Korol;M. Jordan;M. B. Rutzig;A. C. S. Beck

文献摘要

被引文献

相似文献

为了满足延迟和隐私要求,资源匮乏的深度学习应用程序已经迁移到边缘,物联网设备可以将推理处理卸载到本地边缘服务器。由于FPGA已经成功地加速了越来越多的深度学习应用程序(特别是基于CNN的应用程序),因此它们成为边缘平台的有效替代品。然而,边缘应用程序可能会呈现高度不可预测的工作负载,需要在推理处理中具有运行时适应性。虽然一些作品通过在运行时利用不同的修剪率在CPU和GPU平台上应用模型切换,因此推理可以根据一些质量-性能权衡进行调整,但基于FPGA的加速器避免了这种方法,因为它们被合成到特定的CNN模型。在这种情况下,这项工作通过向众所周知的FINN加速器添加额外的适应性级别(即,灵活性),并通过灵活加速器上的快速模型切换(以一些额外逻辑为代价)或通过固定加速器的FPGA重新配置来支持修剪的动态使用。在此基础上,我们开发了AdaFlow:一个框架,它在设计时自动从这些新的可用版本(灵活的和固定的,修剪或不修剪)中构建一个库,该库将在运行时用于根据用户可配置的准确性阈值和当前工作负载条件动态选择给定版本。我们使用两个CNN模型和两个数据集在智能边缘监控应用程序下对AdaFlow进行了评估,结果显示AdaFlow平均处理1.3倍以上的推理,平均提高1.4倍的功率效率,超过最先进的静态部署的超低速加速器。
To meet latency and privacy requirements, resource-hungry deep learning applications have been migrating to the Edge, where IoT devices can offload the inference processing to local Edge servers. Since FPGAs have successfully accelerated an increasing number of deep learning applications (especially CNN-based ones), they emerge as an effective alternative for Edge platforms. However, Edge applications may present highly unpredictable workloads, requiring runtime adaptability in the inference processing. Although some works apply model switching on CPU and GPU platforms by exploiting different pruning rates at runtime, so the inference can adapt according to some quality-performance trade-off, FPGA-based accelerators refrain from this approach since they are synthesized to specific CNN models. In this context, this work enables model switching on FPGAs by adding to the well-known FINN accelerator an extra level of adaptability (i.e., flexibility) and support to the dynamic use of pruning via fast model switch on flexible accelerators, at the cost of some extra logic, or via FPGA reconfigurations of fixed accelerators. From that, we developed AdaFlow: a framework that automatically builds, at design time, a library from these new available versions (flexible and fixed, pruned or not) that will be used, at runtime, to dynamically select a given version according to a user-configurable accuracy threshold and current workload conditions. We have evaluated AdaFlow under a smart Edge surveillance application with two CNN models and two datasets, showing that AdaFlow processes, on average, 1.3× more inferences and increases, on average, 1.4× the power efficiency over state-of-the-art statically deployed dataflow accelerators.