CDS&E: SuperSTARLU - STacked, AcceleRated Algorithms for Sparse Linear Systems
CDS&E: SuperSTARLU - STacked, AcceleRated Algorithms for Sparse Linear Systems
批准号:
1710371
负责人:
Jeffrey Young
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2022-07-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Computing systems and associated software have long had to make trade-offs in terms of performance due to an imbalance between fast processors and their slower memory hierarchies. Newly released technologies for 3D stacked memories provide an opportunity to reduce this imbalance by providing higher memory bandwidths and novel ways for accessing memory. One of the most promising techniques for using 3D stacked memory involves "memory-centric" computation that moves computation as close as possible to main memory. However, there is little understanding of how to best use these new memory technologies in libraries and applications, even as this hardware is slated to be integrated into near-term exascale supercomputing systems. The goal of this project is to understand the techniques and approaches that are needed to fully utilize 3D stacked memories and to demonstrate a useful set of computational primitives that can serve as a template for accelerating large scientific codes with these new memory components.This research considers the specific problem of using "memory-centric" processors effectively to implement sparse primitives as part of a library that supports a number of key scientific applications including radiation transport, fluid flow, and fusion simulations. This library, SuperLU_DIST, is a sparse direct solver library designed for distributed memory multicore systems that has previously been accelerated on both NVIDIA?s graphics co-processors and Intel?s Xeon Phi co-processor. While this prior work has thus far yielded promising speedups, it has also revealed critical and fundamental algorithmic performance bottlenecks related to memory data transfers. This research project will investigate whether these bottlenecks may be mitigated by using emerging memory-centric co-processors. Such co-processors, which include Micron?s Hybrid Memory Cube (HMC) and High-Bandwidth Memory (HBM), combine 3-D stacked memories and FPGAs to provide lower latency, higher bandwidth data transfer, and support for near-memory data processing. The project will use high-level languages like OpenCL to take advantage of such technologies and will utilize a mix of algorithmic advances and software library development to improve application performance. Additionally, this work will lead to a new, open-source release of SuperLU called Super Stacked, Accelerated LU (SuperSTARLU), which will be made available to application developers and will be demonstrated on one of the next-generation systems with memory-centric co-processors, such as NERSC?s Cori.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3457388.3458670
发表时间:
2021
期刊:
The 18th International Conference on Computing Frontiers
影响因子:
--
作者:
[Lavin, Patrick, Young, Jeffrey, Vuduc, Richard, Beard, Jonathan]
通讯作者:
Beard, Jonathan
Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters
多 GPU 集群上大型图的可扩展全对最短路径
DOI:
10.1145/3431379.3460651
发表时间:
2020
期刊:
HPDC '21: Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
作者:
[sao, Piyush, lu, Hao, Kannan, Ramakrishnan, Thakkar, Vijay, Vuduc, Richard, Potok, Thomas]
通讯作者:
Potok, Thomas
A communication-avoiding 3D algorithm for sparse LU factorization on heterogeneous systems
异构系统上稀疏 LU 分解的避免通信 3D 算法
DOI:
10.1016/j.jpdc.2019.03.004
发表时间:
2019
期刊:
Journal of Parallel and Distributed Computing
影响因子:
3.8
作者:
[Sao, Piyush, Li, Xiaoye S., Vuduc, Richard]
通讯作者:
Vuduc, Richard
Performance Implications of NoCs on 3D-Stacked Memories: Insights from the Hybrid Memory Cube
NoC 对 3D 堆叠存储器的性能影响:来自混合存储器立方体的见解
DOI:
10.1109/ispass.2018.00018
发表时间:
2018
期刊:
IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS
影响因子:
--
作者:
[Hadidi, Ramyad, Asgari, Bahar, Young, Jeffrey, Ahmad Mudassar, Burhan, Garg, Kartikay, Krishna, Tushar, Kim, Hyesoon]
通讯作者:
Kim, Hyesoon
DOI:
--
发表时间:
2018
期刊:
IEEE International Parallel and Distributed Processing Symposium Workshops
影响因子:
--
作者:
[Hein, Eric, Conte, Tom, Young, Jeffrey S., Eswar, Srinivas, Li, Jiajia, Lavin, Patrick, Vuduc, Richard, Riedy, Jason]
通讯作者:
Riedy, Jason
共 14 条
CCRI: Medium: Rogues Gallery: A Community Research Infrastructure for Post-Moore Computing
-
批准号:2016701
-
项目类别:Standard Grant
-
资助金额:$135.17万
-
财政年份:2020
-
负责人:Jeffrey Young
-
依托单位:
Conference: Travel Grant for 2002 IEEE-AP-S International Symposium & USNC/USRI National Radio Science Meeting to be held in San Antonio, TX June 16-21, 2002.
-
批准号:0121897
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2002
-
负责人:Jeffrey Young
-
依托单位: