Efficient automatic scheduling of imaging and vision pipelines for the GPU

Efficient automatic scheduling of imaging and vision pipelines for the GPU
复制标题

DOI:
10.1145/3485486
复制
发表时间:
2020-12
影响因子:
--
通讯作者:
Luke Anderson;Andrew Adams;Karima Ma;Tzu-Mao Li;Tian Jin;Jonathan Ragan-Kelley
Luke Anderson;Andrew Adams;Karima Ma;Tzu-Mao Li;Tian Jin;Jonathan Ragan-Kelley
中科院分区:
--
文献类型:
--
作者:
Luke Anderson;Andrew Adams;Karima Ma;Tzu-Mao Li;Tian Jin;Jonathan Ragan-Kelley

文献摘要

相似文献

我们提出了一种新的算法,以快速生成复杂成像和视觉管道的高性能GPU实现,直接从高级HALIDE算法代码中,它是完全自动的。我们取得了成就这首先使用(1)两阶段的搜索算法,该算法首先“冻结”程序的最低成本段的决策,从而使相对较高的时间在重要阶段花费了更多的时间,(2)分组的层次样本策略,以层次的样本来调度基于其结构相似性,然后将我们的样本代表,使我们的成本与少数范围探索(并探索少数示意)(3)发生。通过有效的成本模型将机器学习,程序分析和GPU架构知识进行指导。积极地击败我们的自动结果。
We present a new algorithm to quickly generate high-performance GPU implementations of complex imaging and vision pipelines, directly from high-level Halide algorithm code. It is fully automatic, requiring no schedule templates or hand-optimized kernels. We address the scalability challenge of extending search-based automatic scheduling to map large real-world programs to the deep hierarchies of memory and parallelism on GPU architectures in reasonable compile time. We achieve this using (1) a two-phase search algorithm that first ‘freezes’ decisions for the lowest cost sections of a program, allowing relatively more time to be spent on the important stages, (2) a hierarchical sampling strategy that groups schedules based on their structural similarity, then samples representatives to be evaluated, allowing us to explore a large space with few samples, and (3) memoization of repeated partial schedules, amortizing their cost over all their occurrences. We guide the process with an efficient cost model combining machine learning, program analysis, and GPU architecture knowledge. We evaluate our method’s performance on a diverse suite of real-world imaging and vision pipelines. Our scalability optimizations lead to average compile time speedups of 49x (up to 530x). We find schedules that are on average 1.7x faster than existing automatic solutions (up to 5x), and competitive with what the best human experts were able to achieve in an active effort to beat our automatic results.