ORCHARD: Visual object recognition accelerator based on approximate in-memory processing

ORCHARD: Visual object recognition accelerator based on approximate in-memory processing
复制标题

DOI:
10.1109/iccad.2017.8203756
复制
发表时间:
2017-11
期刊:
2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Yeseong Kim;M. Imani;Tajana Simunic
Yeseong Kim;M. Imani;Tajana Simunic
中科院分区:
其他
文献类型:
--
作者:
Yeseong Kim;M. Imani;Tajana Simunic

文献摘要

被引文献

相似文献

近年来,机器学习在视觉对象识别方面的应用已广泛应用于自动驾驶汽车、健康诊断和家庭自动化等领域。然而,识别过程仍然消耗大量的处理能量,并且产生较高的内存访问数据移动成本。在本文中,我们提出了一种新的硬件加速器设计,称为ORCHARD,它在内存中处理目标识别任务。提出的设计加速了图像特征提取和基于增强的学习算法,这是最先进的图像识别方法的关键子任务。我们通过利用近似计算和新兴的非易失性存储器(NVM)技术来优化识别过程。基于nvm的内存处理允许所提出的设计减少基于cmos的计算开销,极大地提高了系统效率。在我们对电路和设备级模拟进行的评估中,我们表明ORCHARD成功地执行了实际的图像识别任务,包括文本、人脸、行人和车辆识别,通过计算近似获得的精度损失为0.3%。此外,与现有的基于处理器的实现相比,我们的设计显著提高了性能和能源效率,分别提高了376x和1896x。
In recent years, machine learning for visual object recognition has been applied to various domains, e.g., autonomous vehicle, heath diagnose, and home automation. However, the recognition procedures still consume a lot of processing energy and incur a high cost of data movement for memory accesses. In this paper, we propose a novel hardware accelerator design, called ORCHARD, which processes the object recognition tasks inside memory. The proposed design accelerates both the image feature extraction and boosting-based learning algorithm, which are key subtasks of the state-of-the-art image recognition approaches. We optimize the recognition procedures by leveraging approximate computing and emerging non-volatile memory (NVM) technology. The NVM-based in-memory processing allows the proposed design to mitigate the CMOS-based computation overhead, highly improving the system efficiency. In our evaluation conducted on circuit- and device-level simulations, we show that ORCHARD successfully performs practical image recognition tasks, including text, face, pedestrian, and vehicle recognition with 0.3% of accuracy loss made by computation approximation. In addition, our design significantly improves the performance and energy efficiency by up to 376x and 1896x, respectively, compared to the existing processor-based implementation.