AppCiP: Energy-Efficient Approximate Convolution-in-Pixel Scheme for Neural Network Acceleration

AppCiP: Energy-Efficient Approximate Convolution-in-Pixel Scheme for Neural Network Acceleration
复制标题

DOI:
10.1109/jetcas.2023.3242167
复制
发表时间:
2023-03
影响因子:
4.6
通讯作者:
Sepehr Tabrizchi;Ali Nezhadi;Shaahin Angizi;A. Roohi
Sepehr Tabrizchi;Ali Nezhadi;Shaahin Angizi;A. Roohi
中科院分区:
工程技术2区
文献类型:
--
作者:
Sepehr Tabrizchi;Ali Nezhadi;Shaahin Angizi;A. Roohi

文献摘要

被引文献

相似文献

如今,始终在智能和自我的视觉感知系统中始终引起了人们的关注,并被广泛使用。但是,捕获数据并通过后端/云处理器进行分析是能源密集型和长期的,从而导致内存瓶颈和边缘的低速特征提取。本文将AppCIP体系结构作为一种感应和计算集成设计,以有效地在资源有限的传感设备上实现人工智能(AI)。 AppCIP提供了许多独特的功能,包括灰度转换的即时和可重新配置的RGB,高度平行的模拟卷积中的像素,并实现了低精确的Quinary Weighter Wighter网络。这些功能大大减轻了模数转换器和模拟缓冲区的开销,从而大大降低了功耗和开销面积。我们的电路对施加共同模拟结果表明,与考虑不同的CNN工作负载相比,AppCIP的功率消耗效率高约3个数量级。它达到3000的帧速率,效率约为4.12 top/s/w。评估了AppCIP架构在不同数据集上的性能准确性,例如SVHN,PEST,CIFAR-10,MHIST和CBL面部检测,并与最先进的设计进行了比较。获得的结果在其他像素体系结构中/附近的其他处理中表现出最佳结果,而与浮点基线相比,AppCIP的准确性平均降低了不到1%。
Nowadays, always-on intelligent and self-powered visual perception systems have gained considerable attention and are widely used. However, capturing data and analyzing it via a backend/cloud processor are energy-intensive and long-latency, resulting in a memory bottleneck and low-speed feature extraction at the edge. This paper presents AppCiP architecture as a sensing and computing integration design to efficiently enable Artificial Intelligence (AI) on resource-limited sensing devices. AppCiP provides a number of unique capabilities, including instant and reconfigurable RGB to grayscale conversion, highly parallel analog convolution-in-pixel, and realizing low-precision quinary weight neural networks. These features significantly mitigate the overhead of analog-to-digital converters and analog buffers, leading to a considerable reduction in power consumption and area overhead. Our circuit-to-application co-simulation results demonstrate that AppCiP achieves ~3 orders of magnitude higher efficiency on power consumption compared with the fastest existing designs considering different CNN workloads. It reaches a frame rate of 3000 and an efficiency of ~4.12 TOp/s/W. The performance accuracy of the AppCiP architecture on different datasets such as SVHN, Pest, CIFAR-10, MHIST, and CBL Face detection is evaluated and compared with the state-of-the-art design. The obtained results exhibit the best results among other processing in/near pixel architectures, while AppCip only degrades the accuracy by less than 1% on average compared to the floating-point baseline.