Rigel

Rigel
复制标题

DOI:
10.1145/2897824.2925892
复制
发表时间:
2016-07
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
通讯作者:
James Hegarty;Ross G. Daly;Zach DeVito;M. Horowitz;P. Hanrahan;Jonathan Ragan-Kelley
James Hegarty;Ross G. Daly;Zach DeVito;M. Horowitz;P. Hanrahan;Jonathan Ragan-Kelley
中科院分区:
其他
文献类型:
--
作者:
James Hegarty;Ross G. Daly;Zach DeVito;M. Horowitz;P. Hanrahan;Jonathan Ragan-Kelley

文献摘要

被引文献

相似文献

使用定制硬件或 FPGA 实现的图像处理算法的能效和性能比软件高几个数量级。不幸的是,手动将算法转换为适合在这些平台上编译的硬件描述语言通常过于耗时而不实用。最近关于高级图像处理语言的硬件合成的工作表明,可以将模板内核的单速率管道合成到具有可证明的最小缓冲的硬件中。不幸的是,很少有先进的图像处理或视觉算法适合这种高度限制的编程模型。在本文中,我们介绍了 Rigel,它采用我们新的多速率架构中指定的管道并将其降低到 FPGA 实现。我们灵活的多速率架构支持金字塔图像处理、稀疏计算和时空实现权衡。我们演示了立体深度、Lucas-Kanade、SIFT 描述符以及在两个 FPGA 板上运行的高斯金字塔。我们的系统可以合成 FPGA 硬件,吞吐量高达 436 兆像素/秒,运行时间比平板电脑级 ARM CPU 快 297 倍。
Image processing algorithms implemented using custom hardware or FPGAs of can be orders-of-magnitude more energy efficient and performant than software. Unfortunately, converting an algorithm by hand to a hardware description language suitable for compilation on these platforms is frequently too time consuming to be practical. Recent work on hardware synthesis of high-level image processing languages demonstrated that a single-rate pipeline of stencil kernels can be synthesized into hardware with provably minimal buffering. Unfortunately, few advanced image processing or vision algorithms fit into this highly-restricted programming model. In this paper, we present Rigel, which takes pipelines specified in our new multi-rate architecture and lowers them to FPGA implementations. Our flexible multi-rate architecture supports pyramid image processing, sparse computations, and space-time implementation tradeoffs. We demonstrate depth from stereo, Lucas-Kanade, the SIFT descriptor, and a Gaussian pyramid running on two FPGA boards. Our system can synthesize hardware for FPGAs with up to 436 Megapixels/second throughput, and up to 297x faster runtime than a tablet-class ARM CPU.