Reducing overhead in the Uintah framework to support short-lived tasks on GPU-heterogeneous architectures

Reducing overhead in the Uintah framework to support short-lived tasks on GPU-heterogeneous architectures
复制标题

减少 Uintah 框架的开销以支持 GPU 异构架构上的短期任务

DOI:
--
复制
发表时间:
2015
期刊:
WOLFHPC@SC
影响因子:
--
通讯作者:
M. Berzins
M. Berzins
中科院分区:
--
文献类型:
--
作者:
B. Peterson;H. Dasari;A. Humphrey;J. Sutherland;T. Saad;M. Berzins

文献摘要

被引文献

相似文献

Uintah计算框架用于在自适应网格加密网格上使用现代超级计算机并行求解偏微分方程组。Uintah由一个应用层和一个单独的运行时系统构成。Uintah运行时系统基于计算任务的分布式有向无环图(DAG),具有任务调度器,可在CPU核心和节点加速器上高效地调度和执行这些任务。运行时系统识别任务相关性,基于这些相关性在迭代之前创建任务图,为任务准备数据,自动生成MPI消息标签,并在任务计算后管理数据。由于支持更多的内存区域、API调用延迟、内存带宽问题以及开发的额外复杂性,加速器的任务管理比其对应的CPU任务更具挑战性。当任务在几毫秒内计算时,这些挑战是最大的,尤其是那些具有涉及光环数据的基于模板的计算、几乎没有数据重用和/或需要许多计算变量的任务。当前和新兴的异类架构需要在犹他州内解决这些挑战。这项工作不是为了提高现有任务的性能,而是为了减少运行时开销,以允许编写短暂计算任务的开发人员在异类环境中利用Uintah。这项工作分析了在Uintah内管理加速器任务和现有CPU任务的初步方法。这项工作的主要贡献是识别和解决将任务映射到GPU时出现的低效问题,实现新的方案以减少运行时系统开销,引入允许更多任务利用节点上加速器的新功能,并展示这些改进带来的开销减少结果。
The Uintah computational framework is used for the parallel solution of partial differential equations on adaptive mesh refinement grids using modern supercomputers. Uintah is structured with an application layer and a separate runtime system. The Uintah runtime system is based on a distributed directed acyclic graph (DAG) of computational tasks, with a task scheduler that efficiently schedules and execute these tasks on both CPU cores and on-node accelerators. The runtime system identifies task dependencies, creates a taskgraph prior to an iteration based on these dependencies, prepares data for tasks, automatically generates MPI message tags, and manages data after task computation. Managing tasks for accelerators pose significant challenges over their CPU task counterparts due to supporting more memory regions, API call latency, memory bandwidth concerns, and the added complexity of development. These challenges are greatest when tasks compute within a few milliseconds, especially those that have stencil based computations that involve halo data, have little reuse of data, and/or require many computational variables. Current and emerging heterogeneous architectures necessitate addressing these challenges within Uintah. This work is not designed to improve performance of existing tasks, but rather reduce runtime overhead to allow developers writing short-lived computational tasks to utilize Uintah in a heterogeneous environment. This work analyzes an initial approach for managing accelerator tasks alongside existing CPU tasks within Uintah. The principal contribution of this work is to identify and address inefficiencies that arise when mapping tasks onto the GPU, to implement new schemes to reduce runtime system overhead, to introduce new features that allow for more tasks to leverage on-node accelerators, and to show overhead reduction results from these improvements.