课题基金 / 基金详情

SHF: Small: Addressing Challenges for the Next Decade of Massively Parallel NUMA Accelerators

SHF: Small: Addressing Challenges for the Next Decade of Massively Parallel NUMA Accelerators
SHF:小型:应对大规模并行 NUMA 加速器未来十年的挑战
批准号:
1910924
负责人:
Timothy Rogers
金额:
$49.54万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2023-09-30

项目摘要

项目成果

Timothy Rogers的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The physical and economic principles that enabled Dennard scaling and Moore's law in the semiconductor industry have reached their breaking point. However, as the number of transistors economically fabricated on a single chip plateaus, the processor industry has pivoted to create single-package computing systems, composed of multiple sub-components known as chiplets. Chiplets, which communicate via high-bandwidth on-package networks, offer the potential for transparent performance scaling into the next decade. However, chiplets introduce challenging non-uniform memory access characteristics into single-package systems that have traditionally not been subject to these effects. This project develops techniques to overcome the challenges of non-uniform memory accesses on high-performance single- and multi-package systems without programmer intervention. Exploring programmer-transparent scaling mechanisms improves the portability and lifetime of programs, decreasing the cost and complexity of software. Through the creation of course content and undergraduate summer internships, the project fosters an understanding of how to program machines in a post-Moore world and how compute accelerators should be designed to minimize the impact on the end-programmer as system complexity increases.This project develops coordinated data placement and thread scheduling algorithms that leverage static information from the compiler and dynamic information from the runtime system to inform data placement and hardware-based thread scheduling. It advances the state-of-the-art by developing an open-source Graphic Processing Unit (GPU) simulator with a hierarchical interconnect that can be used to model both chiplet-based GPUs and multi-GPU systems. The researchers are exploring compiler informed data placement and thread scheduling in GPUs. Initial results demonstrate that a static analysis of the code can predict the data accessed by GPU threadblocks. Analysis shows that it is possible to determine which threads in a grid share memory pages, and the manner of that sharing, by building new static techniques that add an additional dimension to decades of work on compilers for sequential code. Using static information, in combination with runtime information provided by GPU drivers, the researchers are developing advanced data placement, prefetching, and thread scheduling algorithms. Both future chiplet-based designs and existing multi-GPU systems benefit from the development of these algorithms. Looking beyond the high-bandwidth memory used in GPUs today the project explores the system-level implications of heterogeneous memory in a chiplet-based system. Data placement and thread scheduling have even more importance in GPU systems of the future that make use of high bandwidth memory, traditional dynamic random-access memory, and non-volatile memory. The problem sizes in such systems are anticipated to be so large that opportunistic data placement and thread scheduling are even more critical than in conventional systems. The project uses sharing patterns based on the inter-kernel producer-consumer nature of machine learning workloads to change the program's code layout, runtime data placement, and threadblock scheduling algorithm to maximize locality in multi-node systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/hpca56546.2023.10070957
发表时间: 2023-02
期刊: 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子: --
作者: [Aaron Barnes;Fangjia Shen;Timothy G. Rogers]
通讯作者: Aaron Barnes;Fangjia Shen;Timothy G. Rogers
DOI: 10.1109/micro56248.2022.00040
发表时间: 2022
期刊: 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO
影响因子: --
作者: [Khairy, Mahmoud, Alawneh, Ahmad, Barnes, Aaron, Rogers, Timothy G.]
通讯作者: Rogers, Timothy G.
DOI: 10.1109/micro50266.2020.00086
发表时间: 2020-10
期刊: 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子: --
作者: [Mahmoud Khairy;Vadim Nikiforov;D. Nellans;Timothy G. Rogers]
通讯作者: Mahmoud Khairy;Vadim Nikiforov;D. Nellans;Timothy G. Rogers
DOI: 10.1109/isca45697.2020.00047
发表时间: 2018-10
期刊: 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者: [Mahmoud Khairy;Zhesheng Shen;Tor M. Aamodt;Timothy G. Rogers]
通讯作者: Mahmoud Khairy;Zhesheng Shen;Tor M. Aamodt;Timothy G. Rogers
7
    Autonomous Modelling Solutions for Operational Structural Dynamic Systems
    • 批准号:
      EP/W002140/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $30.94万
    • 财政年份:
      2022
    • 负责人:
      Timothy Rogers
    • 依托单位:
    CAREER: Accessible Accelerators: Leveraging Productive Software on Efficient Hardware
    • 批准号:
      1943379
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $53.13万
    • 财政年份:
      2020
    • 负责人:
      Timothy Rogers
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: