Rhythmic pixel regions: multi-resolution visual sensing system towards high-precision visual computing at low power

Rhythmic pixel regions: multi-resolution visual sensing system towards high-precision visual computing at low power
复制标题

DOI:
10.1145/3445814.3446737
复制
发表时间:
2021-04
期刊:
Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Venkatesh Kodukula;Alexander Shearer;Van Nguyen;Srinivas Lingutla;Yifei Liu;R. Likamwa
Venkatesh Kodukula;Alexander Shearer;Van Nguyen;Srinivas Lingutla;Yifei Liu;R. Likamwa
中科院分区:
其他
文献类型:
--
作者:
Venkatesh Kodukula;Alexander Shearer;Van Nguyen;Srinivas Lingutla;Yifei Liu;R. Likamwa

文献摘要

被引文献

相似文献

高时空分辨率可以为视觉应用提供高精度,这对于捕获视觉特征的细微差别(例如增强现实)特别有用。不幸的是,捕获和处理高时空视觉框架会产生能量昂贵的内存流量。另一方面,低分辨率框架可以减少像素内存吞吐量,但还减少了高精度视觉传感的机会。但是,我们的直觉是,并非需要以统一的分辨率捕获场景的所有部分。在不同区域的图像框架上有选择地和机会降低分辨率可以在节能记忆数据速率下产生高精度的视觉计算。为此,我们开发了一个视觉传感管道体系结构,该架构可以灵活地允许应用程序开发人员动态调整场景中不同“节奏像素区域”的空间分辨率和更新速率。我们开发了一个系统,该系统从商业图像传感器中摄入像素流的标准光栅扫描像素读出模式,但仅在将它们存储在内存中之前对相关的像素进行编码。我们还提出了流硬件,将存储的节奏像素区域流解码为基于传统的框架表示,以输入标准的计算机视觉算法。我们将编码和解码硬件模块集成到现有的视频管道中。最重要的是,我们开发了运行时支持,使开发人员可以灵活地指定区域标签。在三个视觉工作负载上评估我们的系统在Xilinx FPGA平台上显示,界面流量和内存足迹的减少了43-64%,同时提供了可控的任务准确性。
High spatiotemporal resolution can offer high precision for vision applications, which is particularly useful to capture the nuances of visual features, such as for augmented reality. Unfortunately, capturing and processing high spatiotemporal visual frames generates energy-expensive memory traffic. On the other hand, low resolution frames can reduce pixel memory throughput, but reduce also the opportunities of high-precision visual sensing. However, our intuition is that not all parts of the scene need to be captured at a uniform resolution. Selectively and opportunistically reducing resolution for different regions of image frames can yield high-precision visual computing at energy-efficient memory data rates. To this end, we develop a visual sensing pipeline architecture that flexibly allows application developers to dynamically adapt the spatial resolution and update rate of different "rhythmic pixel regions" in the scene. We develop a system that ingests pixel streams from commercial image sensors with their standard raster-scan pixel read-out patterns, but only encodes relevant pixels prior to storing them in the memory. We also present streaming hardware to decode the stored rhythmic pixel region stream into traditional frame-based representations to feed into standard computer vision algorithms. We integrate our encoding and decoding hardware modules into existing video pipelines. On top of this, we develop runtime support allowing developers to flexibly specify the region labels. Evaluating our system on a Xilinx FPGA platform over three vision workloads shows 43-64% reduction in interface traffic and memory footprint, while providing controllable task accuracy.