A scalable computing and memory architecture for variable block size motion estimation on Field-Programmable Gate Arrays

A scalable computing and memory architecture for variable block size motion estimation on Field-Programmable Gate Arrays
复制标题

用于现场可编程门阵列上可变块大小运动估计的可扩展计算和内存架构

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on Field-Programmable Logic and Applications
影响因子:
--
通讯作者:
A. Ye
A. Ye
中科院分区:
--
文献类型:
--
作者:
T. Moorthy;A. Ye

文献摘要

被引文献

相似文献

在本文中,我们研究了使用现场可编程门阵列(FPGA)在设计一个高度可扩展的可变块大小的运动估计架构的H.264/AVC视频编码标准。该架构的可扩展性允许将系统集成到低分辨率视频编码应用的低成本单FPGA解决方案中,以及集成到针对高分辨率应用的高性能多FPGA解决方案中。为了克服FPGA和专用集成电路之间的性能差距,我们的算法智能地增加其并行性的设计规模,同时最大限度地减少内存带宽的使用。该架构的核心计算单元在FPGA上实现,并报告其性能。结果表明,该计算单元能够实现28帧每秒(fps)的性能为640 x480分辨率的VGA视频,而引起的Xilinx XC 5VLX 330 FPGA上只有4%的设备利用率。该架构采用8个计算单元,设备利用率为37%,能够实现31 fps的性能,用于编码完整的1920 x1088逐行HDTV视频。
In this paper, we investigate the use of field-programmable gate arrays (FPGAs) in the design of a highly scalable variable block size motion estimation architecture for the H.264/AVC video encoding standard. The scalability of the architecture allows one to incorporate the system into low cost single FPGA solutions for low-resolution video encoding applications as well as into high performance multi-FPGA solutions targeting high-resolution applications. To overcome the performance gap between FPGAs and application specific integrated circuits, our algorithm intelligently increases its parallelism as the design scales while minimizing the use of memory bandwidth. The core computing unit of the architecture is implemented on FPGAs and its performance is reported. It is shown that the computing unit is able to achieve 28 frames per second (fps) performance for 640x480 resolution VGA video while incurring only 4% device utilization on a Xilinx XC5VLX330 FPGA. With 8 computing units at 37% device utilization, the architecture is able to achieve 31 fps performance for encoding full 1920x1088 progressive HDTV video.