Leo: A Profile-Driven Dynamic Optimization Framework for GPU Applications

Leo: A Profile-Driven Dynamic Optimization Framework for GPU Applications
复制标题

DOI:
--
复制
发表时间:
2014-10
期刊:
--
影响因子:
--
通讯作者:
N. Farooqui;Christopher J. Rossbach;Yuan Yu;K. Schwan
N. Farooqui;Christopher J. Rossbach;Yuan Yu;K. Schwan
中科院分区:
其他
文献类型:
--
作者:
N. Farooqui;Christopher J. Rossbach;Yuan Yu;K. Schwan

文献摘要

被引文献

相似文献

诸如GPU之类的并行体系结构是一种诱人的计算织物,用于渴望性能的开发人员。尽管GPU在许多数据并行应用程序域中启用了速度效果的效果提高,但编写有效的代码实际上可以表现出这些增加的代码是一项非平凡的努力,通常要求开发人员直接在计划模型中直接实现专门的体系结构特征。在GPU上实现良好的性能涉及努力密集型调整,通常要求程序员手动评估多个代码版本,以搜索问题分解与架构 - 和运行时特定于特定时间的参数的最佳组合。对于努力将GPU应用于更通用的计算问题的开发人员,引入不规则的数据结构和访问模式仅适用于加剧这些挑战,并且只会增加所需的努力水平。本文提议使用动态仪器为动态,配置文件驱动的优化提供大部分此类努力。在这个愿景中,程序员使用高级前端编程摘要(例如蒲公英[18])表示应用程序,允许系统而不是程序员探索实现和优化空间。我们认为,这样的系统既可行又急需。我们介绍了这样一个名为Leo的框架的设计。对于一系列基准,我们证明了实施设计的系统可以从1.12到27倍的速度在内核Runtimes中实现,该系统可以改善7-40%的端到端性能。
Parallel architectures like GPUs are a tantalizing compute fabric for performance-hungry developers. While GPUs enable order-of-magnitude performance increases in many data-parallel application domains, writing efficient codes that can actually manifest those increases is a non-trivial endeavor, typically requiring developers to exercise specialized architectural features exposed directly in the programming model. Achieving good performance on GPUs involves effort-intensive tuning, typically requiring the programmer to manually evaluate multiple code versions in search of an optimal combination of problem decomposition with architecture- and runtime-specific parameters. For developers struggling to apply GPUs to more general-purpose computing problems, the introduction of irregular data structures and access patterns serves only to exacerbate these challenges, and only increases the level of effort required. This paper proposes to automate much of this effort using dynamic instrumentation to inform dynamic, profile-driven optimizations. In this vision, the programmer expresses the application using higher-level front-end programming abstractions such as Dandelion [18], allowing the system, rather than the programmer, to explore the implementation and optimization space. We argue that such a system is both feasible and urgently needed. We present the design for such a framework, called Leo. For a range of benchmarks, we demonstrate that a system implementing our design can achieve from 1.12 to 27x speedup in kernel runtimes, which translates to 7-40% improvement for end-to-end performance.