Lightweight asynchronous scheduling in heterogeneous reconfigurable systems

Lightweight asynchronous scheduling in heterogeneous reconfigurable systems
复制标题

异构可重构系统中的轻量级异步调度

DOI:
10.1016/j.sysarc.2022.102398
复制
发表时间:
2022
影响因子:
4.5
通讯作者:
Rodríguez A
Rodríguez A
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rodríguez A

文献摘要

参考文献

被引文献

相似文献

异构嵌入式系统的发展趋势是在同一芯片上集成加速器和通用CPU内核。在这些集成架构中,例如我们在这项工作中针对的Zynq UltraScale+板(CPU+FPGA),加速器和CPU内核之间对共享内存和低开销同步的硬件支持使得探索利用CPU和加速器之间紧密协作的策略成为可能。在本文中,我们提出了一种新的轻量级调度策略,FastFit,针对FPGA加速器,和一个新的调度器,命名为MultiFastFit,异步处理由各种CPU内核和FPGA IP组成的异构系统。与以前最先进的自动调整方法相比,我们的策略显着降低了自动计算接近最优块大小的开销,这使得我们的方法更适合于细粒度应用程序。此外,我们的调度程序MultiFastFit被设计为使计算设备之间的工作的高效协同执行,以这样的方式,所有的设备都是忙碌,同时最大限度地减少负载不平衡。 我们的方法已经使用四个针对低功耗UltraScale+平台精心调整的基准测试进行了评估。我们的实验表明,FastFit策略总是以合理的成本为任何设备配置找到接近最佳的FPGA块大小,即使是细粒度和不规则的应用程序,并且利用所有计算设备的异构CPU+FPGA协同执行通常比仅CPU和仅FPGA执行更快,更节能。我们还将MultiFastFit与其他最先进的调度策略进行了比较,发现它的性能比其他自动调整方法高出2倍,并且在不需要离线搜索理想的CPU-FPGA分区或FPGA块粒度的情况下实现了与手动调整的调度器相似的结果。
The trend for heterogeneous embedded systems is the integration of accelerators and general-purpose CPU cores on the same die. In these integrated architectures, like the Zynq UltraScale+ board (CPU+FPGA) that we target in this work, hardware support for shared memory and low-overhead synchronization between the accelerator and the CPU cores make the case for exploring strategies that exploit a tight collaboration between the CPUs and the accelerator. In this paper we propose a novel lightweight scheduling strategy, FastFit, targeted to FPGA accelerators, and a new scheduler based on it, named MultiFastFit, which asynchronously tackles heterogeneous systems comprised of a variety of CPU cores and FPGA IPs. Our strategy significantly reduces the overhead to automatically compute the near-optimal chunksizes when compared to a previous state-of-the-art auto-tuned approach, which makes our approach more suitable for fine-grained applications. Additionally, our scheduler MultiFastFit has been designed to enable the efficient co-execution of work among compute devices in such a way that all the devices are busy while minimizing the load unbalance. Our approaches have been evaluated using four benchmarks carefully tuned for the low-power UltraScale+ platform. Our experiments demonstrate that the FastFit strategy always finds the near-optimal FPGA chunksize for any device configuration at a reasonable cost, even for fine-grained and irregular applications, and that heterogeneous CPU+FPGA co-executions that exploit all the compute devices are usually faster and more energy efficient than the CPU-only and FPGA-only executions. We have also compared MultiFastFit with other state-of-the-art scheduling strategies, finding that it outperforms other auto-tuned approach up to 2x and it achieves similar results to manually-tuned schedulers without requiring an offline search of the ideal CPU-FPGA partition or FPGA chunk granularity.
异构 Xeon FPGA 平台上的并行多处理和调度
DOI: --
发表时间: 2019
影响因子: 3.3
作者:
Andrés Rodríguez;A. Navarro;R. Asenjo;F. Corbera;R. Gran;D. Suárez;J. Núñez
通讯作者: J. Núñez
一种基于引导自调度的高效消息传递调度器
DOI: --
发表时间: 1989
期刊: International Conference on Supercomputing
影响因子: --
作者:
D. C. Rudolph;C. Polychronopoulos
通讯作者: C. Polychronopoulos
DOI: 10.1109/tcad.2019.2912923
发表时间: 2020-06-01
影响因子: 2.9
作者:
Hosseinabady, Mohammad;Nunez-Yanez, Jose Luis
通讯作者: Nunez-Yanez, Jose Luis
DOI: --
发表时间: 2013
期刊:
影响因子: --
作者:
S. Barrachina;M. Barreda;Sandra Catalán;M. F. Dolz;G. Fabregat;R. Mayo;E. Quintana
通讯作者: E. Quintana