Mapping of H.264 decoding on a multiprocessor architecture

Mapping of H.264 decoding on a multiprocessor architecture
复制标题

DOI:
10.1117/12.476234
复制
发表时间:
2003-05
期刊:
--
影响因子:
--
通讯作者:
Erik B. van der Tol;E. Jaspers;R. Gelderblom
Erik B. van der Tol;E. Jaspers;R. Gelderblom
中科院分区:
其他
文献类型:
--
作者:
Erik B. van der Tol;E. Jaspers;R. Gelderblom

文献摘要

被引文献

相似文献

由于在大批量消费电子产品的竞争领域中,开发成本的重要性日益增加,因此需要通用解决方案来实现设计工作的重用并增加潜在的市场容量。因此,片上系统(SoC)包含越来越多的完全可编程的媒体处理设备,而不是专用系统,后者由于高性能密度而提供了最有吸引力的解决方案。以下是促成这一趋势的原因。首先,SoC越来越多地由其通信基础设施和嵌入式存储器主导,从而使功能单元的成本变得不那么重要。此外,不断增长的设计成本需要可以应用于广泛产品范围的通用解决方案。因此,功能强大的可编程SoC变得越来越有吸引力。然而,为了使功率效率的设计,这也是可扩展的先进的超大规模集成电路技术,并行性应充分利用。任务级并行性和任务级并行性都可以通过例如VLIW多处理器架构来提供。为了提供上述的可扩展性,我们建议在处理器上划分数据,而不是传统的功能分区。这种方法的一个优点是数据的固有局部性,这对于通信高效的软件实现是极其重要的。因此,讨论了软件实现,使得能够例如利用两个处理器架构进行SD分辨率H. 264解码,而高清(HD)解码可以利用执行相同软件的八个处理器系统来实现。实验结果表明,数据通信大大减少高达65%,直接提高了整体性能。除了在内存带宽方面有相当大的改进外,这种新颖的分区概念还提供了一种自然的方法来最佳地平衡所有处理器的负载,从而进一步提高整体加速比。
Due to the increasing significance of development costs in the competitive domain of high-volume consumer electronics, generic solutions are required to enable reuse of the design effort and to increase the potential market volume. As a result from this, Systems-on-Chip (SoCs) contain a growing amount of fully programmable media processing devices as opposed to application-specific systems, which offered the most attractive solutions due to a high performance density. The following motivates this trend. First, SoCs are increasingly dominated by their communication infrastructure and embedded memory, thereby making the cost of the functional units less significant. Moreover, the continuously growing design costs require generic solutions that can be applied over a broad product range. Hence, powerful programmable SoCs are becoming increasingly attractive. However, to enable power-efficient designs, that are also scalable over the advancing VLSI technology, parallelism should be fully exploited. Both task-level and instruction-level parallelism can be provided by means of e.g. a VLIW multiprocessor architecture. To provide the above-mentioned scalability, we propose to partition the data over the processors, instead of traditional functional partitioning. An advantage of this approach is the inherent locality of data, which is extremely important for communication-efficient software implementations. Consequently, a software implementation is discussed, enabling e.g. SD resolution H.264 decoding with a two-processor architecture, whereas High-Definition (HD) decoding can be achieved with an eight-processor system, executing the same software. Experimental results show that the data communication considerably reduces up to 65% directly improving the overall performance. Apart from considerable improvement in memory bandwidth, this novel concept of partitioning offers a natural approach for optimally balancing the load of all processors, thereby further improving the overall speedup.