SACoD: Sensor Algorithm Co-Design Towards Efficient CNN-powered Intelligent PhlatCam

SACoD: Sensor Algorithm Co-Design Towards Efficient CNN-powered Intelligent PhlatCam
复制标题

DOI:
10.1109/iccv48922.2021.00512
复制
发表时间:
2021-10
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Yonggan Fu;Yang Zhang;Yue Wang;Zhihan Lv;Vivek Boominathan;A. Veeraraghavan;Yingyan Lin
Yonggan Fu;Yang Zhang;Yue Wang;Zhihan Lv;Vivek Boominathan;A. Veeraraghavan;Yingyan Lin
中科院分区:
其他
文献类型:
--
作者:
Yonggan Fu;Yang Zhang;Yue Wang;Zhihan Lv;Vivek Boominathan;A. Veeraraghavan;Yingyan Lin

文献摘要

被引文献

相似文献

将卷积神经网络(CNN)驱动的功能集成到物联网(IoT)设备中以实现无处不在的智能“IoT相机”的需求正在蓬勃发展。然而,这种物联网系统的更广泛应用仍然受到两个挑战的限制。首先,一些应用,特别是医疗和可穿戴设备相关的应用,对相机的外形尺寸有严格的要求。其次,强大的CNN通常需要相当大的存储和能源成本,而物联网设备通常资源有限。PhlatCam的外形尺寸可能会降低几个数量级,已成为解决上述第一个挑战的有希望的解决方案,而第二个挑战仍然是一个瓶颈。现有的压缩技术可以解决第二个挑战,但远未实现存储和节能的全部潜力,因为它们主要集中在CNN算法本身。为此,这项工作提出了SACoD,一个传感器算法协同设计框架,以开发更高效的CNN驱动的PhlatCam。特别地,在Phlat-Cam传感器和后端CNN模型中编码的掩码通过差分神经架构搜索在模型参数和架构方面联合优化。广泛的实验,包括模拟和物理测量制造的面具表明,提出的SACoD框架实现了积极的模型压缩和节能,同时保持甚至提高任务的准确性,当基准超过两个国家的最先进的(SOTA)设计与六个数据集在四个不同的视觉任务,包括分类,分割,图像翻译和人脸识别。我们的代码可在https://github.com/RICE-EIC/SACoD上获得。
There has been a booming demand for integrating Convolutional Neural Networks (CNNs) powered functionalities into Internet-of-Thing (IoT) devices to enable ubiquitous intelligent "IoT cameras". However, more extensive applications of such IoT systems are still limited by two challenges. First, some applications, especially medicine-and wearable-related ones, impose stringent requirements on the camera form factor. Second, powerful CNNs often require considerable storage and energy cost, whereas IoT devices often suffer from limited resources. PhlatCam, with its form factor potentially reduced by orders of magnitude, has emerged as a promising solution to the first aforementioned challenge, while the second one remains a bottleneck. Existing compression techniques, which can potentially tackle the second challenge, are far from realizing the full potential in storage and energy reduction, because they mostly focus on the CNN algorithm itself. To this end, this work proposes SACoD, a Sensor Algorithm Co-Design framework to develop more efficient CNN-powered PhlatCam. In particular, the mask coded in the Phlat-Cam sensor and the backend CNN model are jointly optimized in terms of both model parameters and architectures via differential neural architecture search. Extensive experiments including both simulation and physical measurement on manufactured masks show that the proposed SACoD framework achieves aggressive model compression and energy savings while maintaining or even boosting the task accuracy, when benchmarking over two state-of-the-art (SOTA) designs with six datasets across four different vision tasks including classification, segmentation, image translation, and face recognition. Our codes are available at: https://github.com/RICE-EIC/SACoD.