Darkroom: Compiling High-Level Image Processing Code into Hardware Pipelines

Darkroom: Compiling High-Level Image Processing Code into Hardware Pipelines
复制标题

DOI:
10.1145/2601097.2601174
复制
发表时间:
2014-07-01
影响因子:
6.2
通讯作者:
Hanrahan, Pat
Hanrahan, Pat
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hegarty, James;Brunhaver, John;Hanrahan, Pat

文献摘要

被引文献

相似文献

专门的图像信号处理器(ISP)利用图像处理管道的结构使用线条缓冲的架构模式最大程度地减少内存带宽,其中每个阶段之间的所有中间数据都存储在小的片上缓冲区中。这提供了高能效率,可以使用TERA-OP/sec的长管道。电池供电设备中的图像处理,但传统上需要在硬件中进行艰苦的手动设计。基于此模式,我们介绍了暗室,这是一种用于图像处理的语言和编译器。暗室语言的语义使其能够将程序直接编译为线条缓冲管道,其中所有中间值都在本地线条缓冲器存储中,从而消除了与芯片外DRAM的不必要的通信。我们制定了最佳调度线条缓冲管道的问题,以最大程度地减少缓冲区作为整数线性程序。最后,在最佳计划的管道中,Darkroom合成了针对ASIC或FPGA或快速CPU代码的硬件说明。我们评估了一系列应用程序的暗室实现,包括摄像机管道,低级功能检测算法和DeBlurring。对于许多应用程序,我们演示了Gigapixel/sec。在250 MW处的0.5mm(2)低于0.5mm(2)的ASIC硅(在45nm铸造工艺上进行模拟),实时1080p/60视频处理,使用现代FPGA的一小部分资源和数十兆像素/秒。四核X86处理器上的吞吐量。
Specialized image signal processors (ISPs) exploit the structure of image processing pipelines to minimize memory bandwidth using the architectural pattern of line-buffering, where all intermediate data between each stage is stored in small on-chip buffers. This provides high energy efficiency, allowing long pipelines with tera-op/sec. image processing in battery-powered devices, but traditionally requires painstaking manual design in hardware. Based on this pattern, we present Darkroom, a language and compiler for image processing. The semantics of the Darkroom language allow it to compile programs directly into line-buffered pipelines, with all intermediate values in local line-buffer storage, eliminating unnecessary communication with off-chip DRAM. We formulate the problem of optimally scheduling line-buffered pipelines to minimize buffering as an integer linear program. Finally, given an optimally scheduled pipeline, Darkroom synthesizes hardware descriptions for ASIC or FPGA, or fast CPU code. We evaluate Darkroom implementations of a range of applications, including a camera pipeline, low-level feature detection algorithms, and deblurring. For many applications, we demonstrate gigapixel/sec. performance in under 0.5mm(2) of ASIC silicon at 250 mW (simulated on a 45nm foundry process), real-time 1080p/60 video processing using a fraction of the resources of a modern FPGA, and tens of megapixels/sec. of throughput on a quad-core x86 processor.