Scheduling Irregular Dataflow Pipelines on SIMD Architectures

Scheduling Irregular Dataflow Pipelines on SIMD Architectures
复制标题

SIMD 架构上的不规则数据流管道调度

DOI:
10.1145/3380479.3380480
复制
发表时间:
2020
期刊:
WPMVP'20: Proceedings of the 2020 Sixth Workshop on Programming Models for SIMD/Vector Processing
影响因子:
--
通讯作者:
Buhler, Jeremy
Buhler, Jeremy
中科院分区:
--
文献类型:
--
作者:
Plano, Tom;Buhler, Jeremy

文献摘要

参考文献

被引文献

相似文献

流计算通常表现出大量的数据并行性,使它们非常适合SIMD架构。然而,许多这样的计算也表现出不规则性,以数据相关的动态数据速率的形式,这使得高效的SIMD执行具有挑战性。这一挑战的一个方面是需要调度实现为由有限队列连接的阶段的流水线的计算的执行。调度器必须确保高SIMD占用率收集排队项目到vectors和最小化与stages.In这项工作之间切换执行相关的成本,我们提出了AFIE(主动满,非主动空)调度策略不规则流应用程序的SIMD处理器。AFIE可证明地将流水线的每个阶段的输入分组为最小数量的SIMD向量,同时相对于最佳可能策略产生有界数量的开关。这些结果适用于即使不规则性禁止从每个输入到每个阶段产生多少输出的先验知识。我们已经实现了AFIE作为MERCATOR系统的扩展[6],用于在NVIDIA GPU上构建不规则流应用程序。我们描述了如何AFIE调度程序简化MERCATOR的运行时代码和经验衡量新的调度程序的不规则流应用程序的性能提高。
Streaming computations often exhibit substantial data parallelism that makes them well-suited to SIMD architectures. However, many such computations also exhibit irregularity, in the form of data-dependent, dynamic data rates, that makes efficient SIMD execution challenging. One aspect of this challenge is the need to schedule execution of a computation realized as a pipeline of stages connected by finite queues. A scheduler must both ensure high SIMD occupancy by gathering queued items into vectors and minimize costs associated with switching execution between stages.In this work, we present the AFIE (Active Full, Inactive Empty) scheduling policy for irregular streaming applications on SIMD processors. AFIE provably groups inputs to each stage of a pipeline into a minimal number of SIMD vectors while incurring a bounded number of switches relative to the best possible policy. These results apply even though irregularity forbids a priori knowledge of how many outputs will be generated from each input to each stage.We have implemented AFIE as an extension to the MERCATOR system [6] for building irregular streaming applications on NVIDIA GPUs. We describe how the AFIE scheduler simplifies MERCATOR's runtime code and empirically measure the new scheduler's improved performance on irregular streaming applications.
稀疏并行 Delaunay 网格细化
DOI: 10.1145/1248377.1248435
发表时间: 2007
期刊: 2012 2nd IEEE International Conference on Parallel, Distributed and Grid Computing
影响因子: --
作者:
Benoît Hudson;G. Miller;Todd Phillips
通讯作者: Todd Phillips
DOI: 10.1109/cgo.2013.6494989
发表时间: 2013-02
期刊: Proceedings of the 2013 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子: --
作者:
Bin Ren;G. Agrawal;J. Larus;Todd Mytkowicz;T. Poutanen;Wolfram Schulte
通讯作者: Bin Ren;G. Agrawal;J. Larus;Todd Mytkowicz;T. Poutanen;Wolfram Schulte
使用 OpenCL 解决 GPU 架构上的 N-Queens 问题,特别是同步问题
DOI: 10.1109/pdgc.2012.6449926
发表时间: 2012
期刊: 2012 2nd IEEE International Conference on Parallel, Distributed and Grid Computing
影响因子: --
作者:
K. Thouti;S. Sathe
通讯作者: S. Sathe
使用 Auto-Pipe 设计系统将大气切伦科夫望远镜信号处理加速至实时速度
影响因子: 1.4
作者:
Eric J. Tyson;J. Buckley;M. Franklin;R. Chamberlain
通讯作者: R. Chamberlain
DOI: 10.1016/s0022-2836(05)80360-2
发表时间: 1990-10-05
影响因子: 5.6
作者:
ALTSCHUL, SF;GISH, W;LIPMAN, DJ
通讯作者: LIPMAN, DJ