Automatic Irregularity-Aware Fine-Grained Workload Partitioning on Integrated Architectures

Automatic Irregularity-Aware Fine-Grained Workload Partitioning on Integrated Architectures
复制标题

集成架构上的自动不规则感知细粒度工作负载分区

DOI:
10.1109/tkde.2019.2940184
复制
发表时间:
2021-03
影响因子:
8.9
通讯作者:
Du Xiaoyong
Du Xiaoyong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhang Feng;Zhai Jidong;Wu Bo;He Bingsheng;Chen Wenguang;Du Xiaoyong

文献摘要

参考文献

被引文献

相似文献

在同一芯片上同时支持 CPU 和 GPU 的集成架构是一种新兴的、有前景的细粒度 CPU-GPU 协作架构。然而,集成也带来了一些编程和系统优化的挑战,特别是对于图形处理等不规则应用。异构性和不规则性之间复杂的相互作用导致在集成架构上运行不规则应用程序时处理器利用率非常低。此外,CPU 和 GPU 上的细粒度协同处理仍然是一个悬而未决的问题。特别是,在本文中,我们表明之前的 CPU-GPU 协同处理工作负载划分在资源利用率和性能方面远非理想。为了解决这个问题,我们提出了一种名为 FinePar 的系统软件,它考虑了 CPU 和 GPU 的架构差异,并利用集成架构实现的细粒度协作。 FinePar通过不规则感知性能建模和在线自动调优,对不规则工作负载进行分区,实现设备级和线程级负载均衡。我们在两种集成架构上使用图形和稀疏矩阵中的八种不规则应用来评估 FinePar,并将其与最先进的分区方法进行比较。结果表明,FinePar 表现出更好的资源利用率,并且比最佳粗粒度分区方法平均加速 1.6 倍。
The integrated architecture that features both CPU and GPU on the same die is an emerging and promising architecture for fine-grained CPU-GPU collaboration. However, the integration also brings forward several programming and system optimization challenges, especially for irregular applications such as graph processing. The complex interplay between heterogeneity and irregularity leads to very low processor utilization of running irregular applications on integrated architectures. Furthermore, fine-grained co-processing on the CPU and GPU is still an open problem. Particularly, in this paper, we show that the previous workload partitioning for CPU-GPU co-processing is far from ideal in terms of resource utilization and performance. To solve this problem, we propose a system software called FinePar, which considers architectural differences of the CPU and GPU and leverages fine-grained collaboration enabled by integrated architectures. Through irregularity-aware performance modeling and online auto-tuning, FinePar partitions irregular workloads and achieves both device-level and thread-level load balance. We evaluate FinePar with eight irregular applications in graphs and sparse matrices on two integrated architectures and compare it with state-of-the-art partitioning approaches. Results show that FinePar demonstrates better resource utilization and achieves an average of 1.6X speedup over the optimal coarse-grained partitioning method.
DOI: 10.1109/tpds.2013.111
发表时间: 2014-06
影响因子: 5.3
作者:
Jianlong Zhong;Bingsheng He
通讯作者: Jianlong Zhong;Bingsheng He
DOI: 10.14778/3151113.3151122
发表时间: 2017-09
期刊: Proc. VLDB Endow.
影响因子: --
作者:
M. Sha;Yuchen Li;Bingsheng He;K. Tan
通讯作者: M. Sha;Yuchen Li;Bingsheng He;K. Tan
DOI: 10.1007/978-3-642-24449-0_18
发表时间: 2011-09
期刊: --
影响因子: --
作者:
W. Gropp;T. Hoefler;R. Thakur;J. Träff
通讯作者: W. Gropp;T. Hoefler;R. Thakur;J. Träff
DOI: 10.1016/j.procs.2015.05.213
发表时间: 2015
期刊: --
影响因子: --
作者:
A. Vilches;R. Asenjo;A. Navarro;F. Corbera;Rubén Gran Tejero;M. Garzarán
通讯作者: A. Vilches;R. Asenjo;A. Navarro;F. Corbera;Rubén Gran Tejero;M. Garzarán
DOI: --
发表时间: 2010-05
期刊: --
影响因子: --
作者:
J. Ang;Brian W. Barrett;Kyle B. Wheeler;R. Murphy
通讯作者: J. Ang;Brian W. Barrett;Kyle B. Wheeler;R. Murphy