Efficient hierarchical online-autotuning: a case study on polyhedral accelerator mapping

Efficient hierarchical online-autotuning: a case study on polyhedral accelerator mapping
复制标题

DOI:
10.1145/3330345.3330377
复制
发表时间:
2019-06
期刊:
Proceedings of the ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
Philip Pfaffe;T. Grosser;Martin Tillmann
Philip Pfaffe;T. Grosser;Martin Tillmann
中科院分区:
其他
文献类型:
--
作者:
Philip Pfaffe;T. Grosser;Martin Tillmann

文献摘要

被引文献

相似文献

识别优化和并行化编译器应该生成的(接近)最佳程序变体是困难的。自动调优是在可能选项的高维空间中导航的最佳解决方案。然而,为了实用,自动调谐器应该(a)具有高收敛速度并且(B)在面对变化的输入时是鲁棒的。用于离线调谐的当前技术(其中收敛速度不太重要)仅为已知输入提供解决方案,而在线调谐可以是输入敏感的,但当前缺乏收敛速度。在本文中,我们提出了分层在线自调整,一种新的技术,利用结构在搜索空间和潜在的调整问题,以提高收敛速度在线调整。通过对配置中的对称性和冗余性进行建模,并利用领域知识来预测性能,我们将搜索空间的大小减少了几个数量级。将我们的调谐器与用于GPU的多面体并行化编译器相结合,我们表明使用默认参数生成的GEMM GPU内核的性能提高了6倍,并且与OpenTuner相比,调谐过程的收敛速度提高了1.7倍。通过分层调优,我们可以实现始终在线的自动调优部署。
Identifying the (near) optimal program variants an optimizing and parallelizing compiler should generate is known to be difficult. Autotuning is the best solution to navigate the often high-dimensional space of possible options. However, to be practical an autotuner should (a) have high convergence speed and (b) be robust in face of varying inputs. Current techniques for offline tuning, where convergence speed is less important, provide solutions only for known inputs, whereas online tuning can be input sensitive but currently lacks in convergence speed. In this paper, we present hierarchical online-autotuning, a novel technique to exploit structure in the search space and the underlying tuning problem to increase convergence speed during online tuning. By modeling symmetries and redundancies in configurations and by exploiting domain knowledge to predict performance we reduce the search space size by orders of magnitudes. Combining our tuner with a polyhedral parallelizing compiler for GPUs, we show that the performance of a GEMM GPU kernel generated with default parameters is increased by 6× and that the convergence speed of the tuning process is increased by a factor of up to 1.7 compared to OpenTuner. With hierarchical tuning we make the deployment of always-on online-autotuning practical.