Optimizing memory bandwidth exploitation for OpenVX applications on embedded many-core accelerators

Optimizing memory bandwidth exploitation for OpenVX applications on embedded many-core accelerators
复制标题

优化嵌入式多核加速器上 OpenVX 应用程序的内存带宽利用

DOI:
10.1007/s11554-015-0544-0
复制
发表时间:
2018
影响因子:
3
通讯作者:
L. Benini
L. Benini
中科院分区:
计算机科学4区
文献类型:
--
作者:
Giuseppe Tagliavini;Germain Haugou;A. Marongiu;L. Benini

文献摘要

参考文献

被引文献

相似文献

近年来,图像处理已成为移动和嵌入式计算平台的一个关键应用领域。在这种情况下,多核加速器是有效执行高度并行内核的可行解决方案。然而,体系结构的限制对主内存带宽施加了硬限制,并推动了优化复杂多内核应用程序内存使用的软件技术。在这项工作中,我们提出了一套主要基于图形分析和图像平铺的技术,旨在加速以标准OpenVX图形表示的图像处理应用程序在基于集群的多核加速器上的执行。我们已经开发了一个运行时框架,它使用符合OpenVX标准的前端实现了这些技术,并基于OpenCL扩展,可以更显式地控制和有效地重用片上内存,并大大减少了对片外内存的依赖,以存储中间结果。在STHORM多核加速器上进行的实验表明,即使加速器的主存储器带宽受到严重限制,我们的方法也可以大量减少时间和带宽。
In recent years, image processing has been a key application area for mobile and embedded computing platforms. In this context, many-core accelerators are a viable solution to efficiently execute highly parallel kernels. However, architectural constraints impose hard limits on the main memory bandwidth, and push for software techniques which optimize the memory usage of complex multi-kernel applications. In this work, we propose a set of techniques, mainly based on graph analysis and image tiling, targeted to accelerate the execution of image processing applications expressed as standard OpenVX graphs on cluster-based many-core accelerators. We have developed a run-time framework which implements these techniques using a front-end compliant to the OpenVX standard, and based on an OpenCL extension that enables more explicit control and efficient reuse of on-chip memory and greatly reduces the recourse to off-chip memory for storing intermediate results. Experiments performed on the STHORM many-core accelerator demonstrate that our approach leads to massive reduction of time and bandwidth, even when the main memory bandwidth for the accelerator is severely constrained.
DOI: 10.1109/mcse.2010.69
发表时间: 2010-05
影响因子: 2.1
作者:
Stone JE;Gohara D;Shi G
通讯作者: Shi G