A Hardware Architecture for Cell-Based Feature-Extraction and Classification Using Dual-Feature Space

A Hardware Architecture for Cell-Based Feature-Extraction and Classification Using Dual-Feature Space
复制标题

DOI:
10.1109/tcsvt.2017.2726564
复制
发表时间:
2018-10-01
影响因子:
8.4
通讯作者:
Mattausch, Hans Jurgen
Mattausch, Hans Jurgen
中科院分区:
工程技术1区
文献类型:
--
作者:
An, Fengwei;Zhang, Xiangyu;Mattausch, Hans Jurgen

文献摘要

被引文献

相似文献

机器人、移动、可穿戴设备和汽车领域的许多计算机视觉和机器学习应用都受到其实时性能要求的限制。本文报道了一种基于双特征的目标识别协处理器,该协处理器利用了直方图定向梯度(HOG)和haar类描述符以及基于单元格的并行滑动窗口识别机制。HOG和haar描述符的特征提取电路采用基于像素的流水线架构实现,该架构与图像传感器的像素频率同步。在提取每个单元格特征向量后,基于单元格的滑动窗口方案可以对包含该单元格的所有窗口进行并行识别。最近邻搜索分类器分别应用于HOG和haar类特征空间。这两个特征域的互补方面使硬件友好的二进制分类实现了行人检测,并提高了准确性。概念验证原型芯片采用65纳米SOI CMOS,具有薄栅氧化层和埋地氧化层(SOTB CMOS),内核面积为3.22 mm(2),在200 mhz识别工作频率和1 v供电电压下,实现了1.52 nJ/像素的能量效率和30 fps的处理速度,可处理1024 x 1616像素的图像帧。此外,由于基于像素的架构,设计的芯片具有图像大小的灵活性,因此多个芯片可以实现图像缩放。
Many computer-vision and machine-learning applications in robotics, mobile, wearable devices, and automotive domains are constrained by their real-time performance requirements. This paper reports a dual-feature-based object recognition coprocessor that exploits both histogram of oriented gradient (HOG) and Haar-like descriptors with a cell-based parallel sliding-window recognition mechanism. The feature extraction circuitry for HOG and Haar-like descriptors is implemented by a pixel-based pipelined architecture, which synchronizes to the pixel frequency from the image sensor. After extracting each cell feature vector, a cell-based sliding window scheme enables parallelized recognition for all windows, which contain this cell. The nearest neighbor search classifier is, respectively, applied to the HOG and Haar-like feature space. The complementary aspects of the two feature domains enable a hardware-friendly implementation of the binary classification for pedestrian detection with improved accuracy. A proof-ofconcept prototype chip fabricated in a 65-nm SOI CMOS, having thin gate oxide and buried oxide layers (SOTB CMOS), with 3.22-mm(2) core area achieves an energy efficiency of 1.52 nJ/pixel and a processing speed of 30 fps for 1024 x 1616-pixel image frames at 200-MHz recognition working frequency and 1-V supply voltage. Furthermore, multiple chips can implement image scaling, since the designed chip has image-size flexibility attributable to the pixel-based architecture.