A 0.8V Intelligent Vision Sensor with Tiny Convolutional Neural Network and Programmable Weights Using Mixed-Mode Processing-in-Sensor Technique for Image Classification

A 0.8V Intelligent Vision Sensor with Tiny Convolutional Neural Network and Programmable Weights Using Mixed-Mode Processing-in-Sensor Technique for Image Classification
复制标题

具有微型卷积神经网络和可编程权重的 0.8V 智能视觉传感器,使用混合模式传感器内处理技术进行图像分类

DOI:
10.1109/isscc42614.2022.9731675
复制
发表时间:
2022
期刊:
2022 IEEE International Solid- State Circuits Conference (ISSCC)
影响因子:
--
通讯作者:
C. Hsieh
C. Hsieh
中科院分区:
--
文献类型:
--
作者:
Tzu;Guan;Yi;C. Lo;Ren;Meng;K. Tang;C. Hsieh

文献摘要

被引文献

相似文献

用于需要图像分类的应用的具有人工智能(AI)的视觉系统的需求日益增长。然而,成像仪加专用AI加速器解决方案[1]受到成像仪和带有神经网络加速器的配套信号处理器之间的原始图像数据流量造成的功耗和延迟负担的影响,使其不适合低功耗边缘设备中的实时推理。最近,已经开发了具有近传感器或传感器内处理能力的成像器[2]-[6],以提高特定应用的系统效率。在[2]-[4]中,在成像器中实现近传感器Haar类滤波操作以实现人脸检测(FD)。然而,与针对不同任务使用具有可编程权重的卷积神经网络(CNN)不同,此类先前工作的实现功能有限且不可配置。在[5]中,报告了一种具有近传感器模拟乘法累积(MAC)操作的卷积CMOS图像传感器(CIS),用于辅助CNN的第一层计算。然而,卷积CIS对于某些任务来说是不够的,由于层/内核数量的限制,并且需要一个配套的数字加速器来进行所需的操作(整流线性单元:ReLU,最大池化:MP,全连接层:FC等)。一个完整的CNN模型在[6]中,报告了一个模拟卷积CIS,其中包含一个用于CNN实现的5层网络。然而,使用与电容器阵列的电荷共享的模拟MAC操作导致增益损失、低权重分辨率和有限的精度。此外,使用静态赢家通吃电路的ReLU+MP操作是耗电的。为了解决这些问题,我们提出了一种智能视觉传感器(IVS),它具有嵌入式微小CNN模型和可编程权重,可以使用混合模式传感器处理(PIS)技术实现可配置的特征提取和片上图像分类。
Vision systems with artificial intelligence (AI) for applications requiring image classification are in growing demand. However, the imager plus dedicated AI accelerator solution [1] suffers from the burdens of power and latency caused by the raw image data traffic between the imager and the companion signal processor with a neural network accelerator, making it unsuitable for the real-time inference in low-power edge devices. Recently, imagers with near- or in-sensor processing capability have been developed [2]–[6] to improve the system efficiency for specific applications. In [2]–[4], the near-sensor Haar-like filtering operations are implemented in imagers to realize face detection (FD). However, unlike using convolutional neural networks (CNNs) with programmable weights for different tasks, the implemented features of such prior works are limited and not configurable. In [5], a convolutional CMOS image sensor (CIS) with near-sensor analog multiply-accumulate (MAC) operations was reported for assisting with the 1st-layer computations of a CNN. However, the convolutional CIS is inadequate for some tasks, due limits on the numbers of layers/kernels, and needs a companion digital accelerator for the required operations (Rectified Linear Unit: ReLU, Maximum-Pooling: MP, Fully-Connected layer: FC, etc.) of a complete CNN model. In [6], an analog convolutional CIS is reported with a 5-layer network for CNN implementation. However, the analog MAC operations using charge sharing with a capacitor array leads to gain loss, low weight resolution, and limited accuracy. Moreover, the ReLU+MP operation using a static winner-take-all circuit is power hungry. To address these issues, we present an intelligent vision sensor (IVS) with an embedded tiny CNN model and programmable weights to achieve configurable feature extraction and on-chip image classification using a mixed-mode processing-in-sensor (PIS) technique.