Optimizing memory bandwidth exploitation for OpenVX applications on embedded many-core accelerators
Optimizing memory bandwidth exploitation for OpenVX applications on embedded many-core accelerators
复制标题
优化嵌入式多核加速器上 OpenVX 应用程序的内存带宽利用
DOI:
10.1007/s11554-015-0544-0
复制
发表时间:
2018
影响因子:
3
通讯作者:
L. Benini
中科院分区:
文献类型:
--
作者:
Giuseppe Tagliavini;Germain Haugou;A. Marongiu;L. Benini
In recent years, image processing has been a key application area for mobile and embedded computing platforms. In this context, many-core accelerators are a viable solution to efficiently execute highly parallel kernels. However, architectural constraints impose hard limits on the main memory bandwidth, and push for software techniques which optimize the memory usage of complex multi-kernel applications. In this work, we propose a set of techniques, mainly based on graph analysis and image tiling, targeted to accelerate the execution of image processing applications expressed as standard OpenVX graphs on cluster-based many-core accelerators. We have developed a run-time framework which implements these techniques using a front-end compliant to the OpenVX standard, and based on an OpenCL extension that enables more explicit control and efficient reuse of on-chip memory and greatly reduces the recourse to off-chip memory for storing intermediate results. Experiments performed on the STHORM many-core accelerator demonstrate that our approach leads to massive reduction of time and bandwidth, even when the main memory bandwidth for the accelerator is severely constrained.
影响因子:
2.1
作者:
Stone JE;Gohara D;Shi G
通讯作者:
Shi G