A Real-Time 17-Scale Object Detection Accelerator With Adaptive 2000-Stage Classification in 65 nm CMOS

A Real-Time 17-Scale Object Detection Accelerator With Adaptive 2000-Stage Classification in 65 nm CMOS
复制标题

DOI:
10.1109/tcsi.2019.2921714
复制
发表时间:
2019-06
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
Minkyu Kim;Abinash Mohanty;Deepak Kadetotad;Luning Wei;Xiaofei He;Yu Cao;Jae-sun Seo
Minkyu Kim;Abinash Mohanty;Deepak Kadetotad;Luning Wei;Xiaofei He;Yu Cao;Jae-sun Seo
中科院分区:
其他
文献类型:
--
作者:
Minkyu Kim;Abinash Mohanty;Deepak Kadetotad;Luning Wei;Xiaofei He;Yu Cao;Jae-sun Seo

文献摘要

相似文献

机器学习已经在包括对象检测、图像/视频分类和自然语言处理在内的应用中变得无处不在。虽然机器学习算法已经成功地应用于许多实际应用中,但这些算法的准确、快速和低功耗硬件实现仍然是一项具有挑战性的任务,特别是对于物联网(IoT)、自动驾驶汽车和智能无人机等移动的系统。本文提出了一种用于目标检测的节能可编程ASIC加速器。我们的ASIC加速器支持多类(例如,人脸、交通标志、汽车牌照和行人),可编程,一张图像中有多个对象(多达50个),具有不同的尺寸(17级支持,6级缩小/11级放大),以及高精度(FDDB/AFW/BTSD/Caltech数据集的AP为0.87/0.81/0.72/0.76)。我们设计了一个具有2,000个分类器的积分通道检测器,用于刚性提升模板,其中用于分类的级数可以根据搜索窗口的内容自适应地控制。与支持向量机(SVM)和可变形零件模型(DEPELLED)设计相比,这可以用更模块化的硬件来实现。通过联合优化算法和高效的硬件架构,在65 nm CMOS实现的原型芯片演示了20-50帧/秒的实时目标检测,低功耗为22.5-181.7 mW(0.54-1.75 nJ/像素),0.58-1.1 V电源。
Machine learning has become ubiquitous in applications including object detection, image/video classification, and natural language processing. While machine learning algorithms have been successfully used in many practical applications, accurate, fast, and low-power hardware implementations of such algorithms is still a challenging task, especially for mobile systems such as Internet of Things (IoT), autonomous vehicles, and smart drones. This paper presents an energy-efficient programmable ASIC accelerator for object detection. Our ASIC accelerator supports multi-class (e.g., face, traffic sign, car license plate, and pedestrian) that are programmable, many-object (up to 50) in one image with different sizes (17-scale support with 6 down-/11 up-scaling), and high accuracy (AP of 0.87/0.81/0.72/0.76 for FDDB/AFW/BTSD/Caltech datasets). We designed an integral channel detector with 2,000 classifiers for rigid boosted templates, where the number of stages used for classification can be adaptively controlled depending on the content of the search window. This can be implemented with a more modular hardware, compared to support vector machine (SVM) and deformable parts model (DPM) designs. By jointly optimizing the algorithm and the efficient hardware architecture, the prototype chip implemented in 65nm CMOS demonstrates real-time object detection of 20–50 frames/s with low power consumption of 22.5–181.7 mW (0.54–1.75 nJ/pixel) at 0.58–1.1 V supply.