课题基金 / 基金详情

SHF: Small: Locality Aware Scheduling in Multi-GPU Systems

SHF: Small: Locality Aware Scheduling in Multi-GPU Systems
SHF:小型:多 GPU 系统中的局部感知调度
批准号:
1907401
负责人:
Laxmi Bhuyan
金额:
$43.16万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2024-09-30

项目摘要

项目成果

Laxmi Bhuyan的其他基金

相似基金

相关文献

中文摘要
翻译
由中央处理单元(CPU)和图形处理单元(GPU)组成的异类多处理器体系结构越来越多地用于加速高性能计算(HPC)和云计算等并行工作负载。与传统的多核CPU相比,GPU在性能上有了显著的提高,因此被广泛用作加速器。采用多个GPU进一步加快执行速度,提高存储容量。当前的多GPU架构,如DGX,提供超高带宽的NVLink通信,以便在GPU之间直接传输数据。然而,基于不同的内存和通信模型将这些计算和数据划分到多个GPU中对程序员来说是一个巨大的挑战。该项目为不同的应用程序开发了基于图形的分区技术,考虑到GPU中计算的数据局部性。其次,目前关于异构性调度的文献没有考虑GPU内部的处理,将其留给了制造商。该项目还开发了一个基于局部性的线程块(TB)调度器,将相同的基于图的技术扩展到缓存块共享。首先,它开发了用于测量在多GPU架构中执行的计算和通信成本的微观基准。开发了一个评测工具来测量TB之间的数据共享程度,以供GPU执行。其次,为多GPU数据划分设计了邻接图,其中顶点表示计算,边表示顶点之间的通信代价。对于GPU内部的数据共享,也开发了一个类似的图模型,其中顶点表示TB,边表示TB之间的共享块数量。第三,利用已知的启发式算法和软件,提出了一种邻接图的递归双划分技术,以实现多GPU系统中各分区之间的负载均衡,并最小化各分区之间的通信开销。考虑到二级缓存的大小和GPU内部的资源限制,提出了TB级调度。第四,将该技术扩展到在异构型多处理器中的CPU和GPU之间划分数据和计算。最后,分析了LU分解和Wavefront这两个常规应用,并利用GPU体系结构进行了实际实现,实现了多GPU调度。还分析了Rodinia和CUDA-SDK基准中的一些非常规应用程序,以开发图形模型并在GPGPUSim上执行它们,以验证GPU内的结核病调度。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Heterogeneous multiprocessor architectures consisting of Central Processing Units (CPUs) and Graphical Processing Units (GPUs) are increasingly used to accelerate parallel workloads like High Performance Computing (HPC) and cloud computing. GPUs provide significant improvements in performance compared to traditional multi-core CPUs, and therefore, are heavily used as accelerators. Multiple GPUs are employed to further speed up the execution and improve storage capacity. Current multi-GPU architectures, such as DGX, provide ultra-high bandwidth NVLink communication to transfer the data directly between the GPUs. However, partitioning those computations and data in multi-GPUs based on various memory and communication models poses a tremendous challenge to the programmers. This project develops graph-based partitioning techniques for different applications considering data locality among the computations in the GPUs. Secondly, the current literature on heterogeneous scheduling does not consider processing inside the GPU, leaving it to the manufacturer. This project also develops a locality-based Thread Block (TB) scheduler by extending the same graph-based technique to cache block sharing.The project is carried out in several steps. First, it develops micro-benchmarks for measuring the computation and communication cost for execution in a multi-GPU architecture. A profiling tool is developed to measure the data sharing among the TBs for GPU execution. Second, an adjacency graph is designed for the multi-GPU data partition, where the vertices represent the computation, and edges represent the communication cost between the vertices. A similar graph model is also developed for data sharing inside a GPU, where vertices represent the TBs and edges represent the number of shared blocks between the TBs. Third, a recursive bi-partitioning technique is developed for the adjacency graph using known heuristics and software to achieve load balance among the partitions and minimize the communication cost between the partitions in a multi-GPU system. TB scheduling is also proposed considering the L2 cache size and the resource limit inside a GPU. Fourth, the technique is extended to partition data and computations between CPUs and GPUs in a heterogeneous multiprocessor. Finally, two regular applications, LU decomposition and Wavefront, are analyzed, and multi-GPU scheduling is developed through real implementation using GPU architectures. Some irregular applications from the Rodinia and CUDA-SDK benchmarks are also analyzed to develop graph models and execute them on the GPGPUSim for verification of the TB scheduling inside the GPU.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
GreenMD: Energy-efficient Matrix Decomposition on Heterogeneous Multi-GPU Systems
GreenMD:异构多 GPU 系统上的节能矩阵分解
DOI: 10.1145/3583590
发表时间: 2023
期刊: ACM Transactions on Parallel Computing
影响因子: 1.6
作者: [Zamani, Hadi, Bhuyan, Laxmi, Chen, Jieyang, Chen, Zizhong]
通讯作者: Chen, Zizhong
DOI: 10.1145/3572848.3577496
发表时间: 2023-01
期刊: Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming
影响因子: --
作者: [Jieyang Chen;Xin Liang;Kai Zhao;H. Sabzi;L. Bhuyan;Zizhong Chen]
通讯作者: Jieyang Chen;Xin Liang;Kai Zhao;H. Sabzi;L. Bhuyan;Zizhong Chen
Travel: Student Travel Support to NAS 2021
  • 批准号:
    2139217
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2021
  • 负责人:
    Laxmi Bhuyan
  • 依托单位:
SHF: Medium: Energy Efficient Computing on GPU-based Heterogeneous Systems
  • 批准号:
    1513201
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $75.0万
  • 财政年份:
    2015
  • 负责人:
    Laxmi Bhuyan
  • 依托单位:
SHF: Small: Efficient CPU-GPU Communication for Heterogeneous Architectures
  • 批准号:
    1423108
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.9万
  • 财政年份:
    2014
  • 负责人:
    Laxmi Bhuyan
  • 依托单位:
EAGER: Developing a Programming Environment for Heterogenous Multiprocessors
  • 批准号:
    1157377
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.93万
  • 财政年份:
    2012
  • 负责人:
    Laxmi Bhuyan
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: