课题基金 / 基金详情

CDS&E: SuperSTARLU - STacked, AcceleRated Algorithms for Sparse Linear Systems

CDS&E: SuperSTARLU - STacked, AcceleRated Algorithms for Sparse Linear Systems
CDS
批准号:
1710371
负责人:
Jeffrey Young
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2022-07-31

项目摘要

项目成果

Jeffrey Young的其他基金

相关文献

中文摘要
翻译
由于快速处理器与其较慢的存储器层次结构之间的不平衡,计算系统和相关联的软件长期以来不得不在性能方面进行权衡。新发布的3D堆叠存储器技术提供了一个机会,通过提供更高的存储器带宽和访问存储器的新方法来减少这种不平衡。使用3D堆栈内存的最有前途的技术之一涉及“以内存为中心”的计算,该计算将计算尽可能靠近主内存。然而,人们对如何在库和应用程序中最好地使用这些新的内存技术知之甚少,即使这种硬件预计将被集成到近期的百亿亿次超级计算系统中。本项目的目标是了解充分利用3D堆叠存储器所需的技术和方法,并展示一组有用的计算原语,这些计算原语可以作为使用这些新的存储器组件加速大型科学代码的模板。处理器有效地实现稀疏基元作为一个库的一部分,支持一些关键的科学应用,包括辐射传输,流体流动和融合模拟。这个库,SuperLU_DIST,是一个稀疏的直接求解器库,专为分布式内存多核系统,以前已经在NVIDIA?的图形协处理器和英特尔?的Xeon Phi协处理器。 虽然这项先前的工作迄今为止已经产生了有希望的加速,但它也揭示了与内存数据传输相关的关键和基本的算法性能瓶颈。这个研究项目将调查这些瓶颈是否可以通过使用新兴的以内存为中心的协处理器来缓解。这种协处理器,其中包括微米?的混合存储器立方体(HMC)和高带宽存储器(HBM),将联合收割机3-D堆栈存储器和FPGA结合起来,提供更低的延迟、更高的带宽数据传输,并支持近存储器数据处理。该项目将使用OpenCL等高级语言来利用这些技术,并将利用算法进步和软件库开发的组合来提高应用程序的性能。此外,这项工作将导致一个新的,开源版本的超级LU称为超级堆叠,加速LU(SuperSTARLU),这将提供给应用程序开发人员,并将在下一代系统与内存为中心的协处理器,如NERSC?是科里。
英文摘要
Computing systems and associated software have long had to make trade-offs in terms of performance due to an imbalance between fast processors and their slower memory hierarchies. Newly released technologies for 3D stacked memories provide an opportunity to reduce this imbalance by providing higher memory bandwidths and novel ways for accessing memory. One of the most promising techniques for using 3D stacked memory involves "memory-centric" computation that moves computation as close as possible to main memory. However, there is little understanding of how to best use these new memory technologies in libraries and applications, even as this hardware is slated to be integrated into near-term exascale supercomputing systems. The goal of this project is to understand the techniques and approaches that are needed to fully utilize 3D stacked memories and to demonstrate a useful set of computational primitives that can serve as a template for accelerating large scientific codes with these new memory components.This research considers the specific problem of using "memory-centric" processors effectively to implement sparse primitives as part of a library that supports a number of key scientific applications including radiation transport, fluid flow, and fusion simulations. This library, SuperLU_DIST, is a sparse direct solver library designed for distributed memory multicore systems that has previously been accelerated on both NVIDIA?s graphics co-processors and Intel?s Xeon Phi co-processor. While this prior work has thus far yielded promising speedups, it has also revealed critical and fundamental algorithmic performance bottlenecks related to memory data transfers. This research project will investigate whether these bottlenecks may be mitigated by using emerging memory-centric co-processors. Such co-processors, which include Micron?s Hybrid Memory Cube (HMC) and High-Bandwidth Memory (HBM), combine 3-D stacked memories and FPGAs to provide lower latency, higher bandwidth data transfer, and support for near-memory data processing. The project will use high-level languages like OpenCL to take advantage of such technologies and will utilize a mix of algorithmic advances and software library development to improve application performance. Additionally, this work will lead to a new, open-source release of SuperLU called Super Stacked, Accelerated LU (SuperSTARLU), which will be made available to application developers and will be demonstrated on one of the next-generation systems with memory-centric co-processors, such as NERSC?s Cori.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
Online model swapping for architectural simulation
建筑模拟的在线模型交换
DOI: 10.1145/3457388.3458670
发表时间: 2021
期刊: The 18th International Conference on Computing Frontiers
影响因子: --
作者: [Lavin, Patrick, Young, Jeffrey, Vuduc, Richard, Beard, Jonathan]
通讯作者: Beard, Jonathan
Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters
多 GPU 集群上大型图的可扩展全对最短路径
DOI: 10.1145/3431379.3460651
发表时间: 2020
期刊: HPDC '21: Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing
影响因子: --
作者: [sao, Piyush, lu, Hao, Kannan, Ramakrishnan, Thakkar, Vijay, Vuduc, Richard, Potok, Thomas]
通讯作者: Potok, Thomas
DOI: 10.1016/j.jpdc.2019.03.004
发表时间: 2019
期刊: Journal of Parallel and Distributed Computing
影响因子: 3.8
作者: [Sao, Piyush, Li, Xiaoye S., Vuduc, Richard]
通讯作者: Vuduc, Richard
Performance Implications of NoCs on 3D-Stacked Memories: Insights from the Hybrid Memory Cube
NoC 对 3D 堆叠存储器的性能影响:来自混合存储器立方体的见解
DOI: 10.1109/ispass.2018.00018
发表时间: 2018
期刊: IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS
影响因子: --
作者: [Hadidi, Ramyad, Asgari, Bahar, Young, Jeffrey, Ahmad Mudassar, Burhan, Garg, Kartikay, Krishna, Tushar, Kim, Hyesoon]
通讯作者: Kim, Hyesoon
共 14 条
    CCRI: Medium: Rogues Gallery: A Community Research Infrastructure for Post-Moore Computing
    • 批准号:
      2016701
    • 项目类别:
      Standard Grant
    • 资助金额:
      $135.17万
    • 财政年份:
      2020
    • 负责人:
      Jeffrey Young
    • 依托单位: