System-Level Design and Integration of a Prototype AR/VR Hardware Featuring a Custom Low-Power DNN Accelerator Chip in 7nm Technology for Codec Avatars

System-Level Design and Integration of a Prototype AR/VR Hardware Featuring a Custom Low-Power DNN Accelerator Chip in 7nm Technology for Codec Avatars
复制标题

原型 AR/VR 硬件的系统级设计和集成,采用 7nm 技术的定制低功耗 DNN 加速器芯片,用于编解码器化身

DOI:
--
复制
发表时间:
2022
期刊:
IEEE Custom Integrated Circuits Conference
影响因子:
--
通讯作者:
E. Beigné
E. Beigné
中科院分区:
--
文献类型:
--
作者:
H. Sumbul;Tony F. Wu;Yuecheng Li;Syed Shakib Sarwar;W. Koven;Eli Murphy;Xingxing Cai;E. Ansari;D. Morris;Huichu Liu;Doyun Kim;E. Beigné

文献摘要

被引文献

相似文献

增强现实/虚拟现实(AR/VR)设备旨在将Metverse中的人们与照片级真实感虚拟化身(Codec Avatars)联系起来。然而,由于AR/VR设备的功率和外形限制有限,为编解码器阿凡达工作负载提供高视觉性能对移动SoC来说是一项具有挑战性的任务。设备上、本地、接近传感器的处理可提供最佳的系统级能效,并可长期实现强大的安全和隐私功能。在这项工作中,我们提出了一个定制的,小规模的移动SoC原型,它实现了对Codec阿凡达模型的跑动眼睛凝视提取的节能性能。该测试芯片采用7 nm工艺节点制造,具有神经网络(NN)加速器,由1024乘法累加(MAC)阵列、2MB片上SRAM和32位RISC-V CPU组成。这款特色测试芯片被集成在原型移动VR头戴式耳机上,以运行Codec阿凡达应用程序。这项工作旨在展示系统级集成、硬件感知模型定制和电路级加速的全栈设计考虑因素,以满足编解码器阿凡达演示中具有挑战性的移动AR/VR SoC规范。通过重新设计基于卷积神经网络(CNN)的眼睛注视提取模型,并针对硬件进行定制,整个模型适合芯片,以降低系统级能量和片外存储器访问的延迟成本。通过在电路级高效地加速卷积运算,所提出的原型SoC在低外形因数下以低功耗实现了每秒30帧的性能。采用本文提出的全栈设计思想,单幅传感器图像从输入到输出的时间为16.5ms,整个CNN模型的功耗为22.7 mW。结果,该测试芯片在2.56 mm2的硅片面积内实现了375微焦耳/帧/眼的能效。
Augmented Reality / Virtual Reality (AR/VR) devices aim to connect people in the Metaverse with photorealistic virtual avatars, referred to as “Codec Avatars”. Delivering a high visual performance for Codec Avatar workloads, however, is a challenging task for mobile SoCs as AR/VR devices have limited power and form factor constraints. On-device, local, near-sensor processing provides the best system-level energy-efficiency and enables strong security and privacy features in the long run. In this work, we present a custom-built, prototype small-scale mobile SoC that achieves energy-efficient performance for running eye gaze extraction of the Codec Avatar model. The test-chip, fabricated in 7nm technology node, features a Neural Network (NN) accelerator consisting of a 1024 Multiply-Accumulate (MAC) array, 2MB on-chip SRAM, and a 32bit RISC-V CPU. The featured test-chip is integrated on a prototype mobile VR headset to run the Codec Avatar application. This work aims to show the full stack design considerations of system-level integration, hardware-aware model customization, and circuit-level acceleration to meet the challenging mobile AR/VR SoC specifications for a Codec Avatar demonstration. By re-architecting the Convolutional NN (CNN) based eye gaze extraction model and tailoring it for the hardware, the entire model fits on the chip to mitigate system-level energy and latency cost of off-chip memory accesses. By efficiently accelerating the convolution operation at the circuit-level, the presented prototype SoC achieves 30 frames per second performance with low-power consumption at low form factors. With the full-stack design considerations presented in this work, the featured test-chip consumes 22.7mW power to run inference on the entire CNN model in 16.5ms from input to output for a single sensor image. As a result, the test-chip achieves 375 µJ/frame/eye energy-efficiency within a 2.56 mm2 silicon area.