Implementation of stereo matching using a high level compiler for parallel computing acceleration

Implementation of stereo matching using a high level compiler for parallel computing acceleration
复制标题

DOI:
10.1145/2425836.2425892
复制
发表时间:
2012-11
期刊:
--
影响因子:
--
通讯作者:
Jinglin Zhang;J. Nezan;J.-G. Cousin;E. Raffin
Jinglin Zhang;J. Nezan;J.-G. Cousin;E. Raffin
中科院分区:
其他
文献类型:
--
作者:
Jinglin Zhang;J. Nezan;J.-G. Cousin;E. Raffin

文献摘要

被引文献

相似文献

在通用计算的许多领域,异质计算系统通过CPU、GPU等加速器提高了并行计算的性能。随着硬件的发展,计算统一设备架构(CUDA)和开放计算语言(OpenCL)等软件开发试图为并行计算提供一个简单而可视化的框架。但事实证明,这比在CPU平台上编程更难进行性能优化。对于一种并行计算应用,不同的硬件平台有不同的配置和参数。本文应用混合多核并行编程技术(HMPP)自动生成适用于GPU平台的可调代码,给出了立体匹配的实现结果,并与C代码版本和手动CUDA版本进行了详细的比较。实验结果表明,与CUDA实现相比,默认和优化后的HMPP具有大致相同的性能和更好的视差图质量。
Heterogeneous computing systems increase the performance of parallel computing in many domains of general purpose computing with CPU, GPU and other accelerators. With Hardware developments, the software developments like Compute Unified Device Architecture (CUDA) and Open Computing Language (OpenCL) try to offer a simple and visual framework for parallel computing. But it turns out to be more difficult than programming on CPU platform for optimization of performance. For one kind of parallel computing application, there are different configurations and parameters for various hardware platforms. In this paper, we apply the Hybrid Multi-cores Parallel Programming (HMPP) to automatically generate tunable code for GPU platform and show the results of implementation of Stereo Matching with detailed comparison with C code version and manual CUDA version. The experimental results show that default and optimized HMPP have approximately the same performance and the better quality of disparity map compared with CUDA implementation.